Stanford NLP Group

@stanfordnlp.bsky.social

Computational Linguists—Natural Language—Machine Learning

Life update! Excited to announce that I’ll be starting as an assistant professor at Cornell Info Sci in August 2026! I’ll be recruiting students this upcoming cycle! An abundance of thanks to all my mentors and friends who helped make this possible!!

The Model Context Protocol is cool because it gives external developers a way to add meaningful functionality on top of LLM platforms. To limit test this, I made a "Realtime Voice" MCP using free STT, VAD, and TTS systems. The result is a janky, but makes me me excited about the ecosystem to come!

Stanford scholars introduced an open-source AI agent that learns how to navigate websites by mimicking childhood learning – an approach that could lead to more efficient, transparent, and privacy-conscious AI: hai.stanford.edu/news/an-open... @chrmanning.bsky.social @shikharmurty.bsky.social

An Open-Source AI Agent for Doing Tasks on the Web | Stanford HAI

NNetNav learns how to navigate websites by mimicking childhood learning through exploration.

hai.stanford.edu

Check it out for cool plots like this about how affinities between words in sentences and how they can show how Green Day isn't like green paint or green tea. And congrats to @coryshain.bsky.social and the CLiMB lab! climblab.org

Bild
Cory Shain@coryshain.bsky.social · last yr.

🚨 First preprint from the lab! 🚨 Josh Rozner (w/@weissweiler.bsky.social and @kmahowald.bsky.social) uses counterfactual experiments on LMs to show that word distributions can provide a learning signal for diverse syntactic constructions, including some hard cases.

🚨 First preprint from the lab! 🚨 Josh Rozner (w/@weissweiler.bsky.social and @kmahowald.bsky.social) uses counterfactual experiments on LMs to show that word distributions can provide a learning signal for diverse syntactic constructions, including some hard cases.

Constructions are Revealed in Word Distributions

Construction grammar posits that constructions (form-meaning pairings) are acquired through experience with language (the distributional learning hypothesis). But how much information about…

arxiv.org

I am concerned about AI but late at night, alone working on a proposal, I was glad ChatGPT had my back as I hit submit 😀.. Reminded me of @chrmanning.bsky.social’s mention in a talk of the 'Real World Utility Test' - early adoption of tech moves forward when it’s genuinely useful, concerns and all.

We are getting closer to have agents operating in the real physical world. However, can we trust frontier models to make embodied decisions 🎮 aligned with human norms 👩‍⚖️ ? With EgoNormia, a 1.8k ego-centric video 🥽 QA benchmark, we show that this is surprisingly challenging!

1/13 New Paper!! We try to understand why some LMs self-improve their reasoning while others hit a wall. The key? Cognitive behaviors! Read our paper on how the right cognitive behaviors can make all the difference in a model's ability to improve with RL! 🧵

Bild

In 2013, at AKBC 2013 and other workshops, I gave a talk titled “Texts are Knowledge”. This was well before there were any transformer LLMs—indeed before the invention of attention—and my early neural NLP ideas were rudimentary. 🔮 Nevertheless, the talk was quite prophetic!

BildBildBildBild

🧵Introducing LangProBe: the first benchmark testing where and how composing LLMs into language programs affects cost-quality tradeoffs! We find that, on avg across diverse tasks, smaller models within optimized programs beat calls to larger models at a fraction of the cost.

Bild

AI won’t reshape education without tackling real problems. Why are we not visiting schools or talking to teachers? A year ago, I partnered with a district facing a major challenge. Instead of doing AI x Education research in isolation, I focused on their real needs.🧵

Bild

We’ve been thrilled by the positive reception to Gemini 2.0 Flash Thinking we discussed in December. Today we’re sharing an experimental update w/improved performance on math, science, and multimodal reasoning benchmarks 📈: • AIME: 73.3% • GPQA: 74.2% • MMMU: 75.4%

Bild

LM agents today primarily aim to automate tasks. Can we turn them into collaborative teammates? 🤖➕👤 Introducing Collaborative Gym (Co-Gym), a framework for enabling & evaluating human-agent collaboration! I now get used to agents proactively seeking confirmations or my deep thinking.(🧵 with video)

Bild

The lawyers are really bringing it this morning at the IMLS National Forum on Data Speculations and I love it. "There is no principled difference between the fair use argument for text and data mining and the fair use argument for AI. If AI is theft, so is your scholarship." 🔥

postdoc opportunity in @alexwoolgar.bsky.social and my lab, based in Cambridge UK! seeking someone with excellent analytical skills to join our project using time-resolved human neuroimaging to study receptive language processing in non-speaking autistic individuals 🧠✨ www.jobs.cam.ac.uk/job/48835/

Postdoctoral Research Associate (Fixed Term) - Job Opportunities - University of Cambridge

Postdoctoral Research Associate (Fixed Term) in the MRC Cognition and Brain Sciences Unit at the University of Cambridge.

jobs.cam.ac.uk

"Mission: Impossible" was featured in Quanta Magazine! Big thank you to @benbenbrubaker.bsky.social for the wonderful article covering our work on impossible languages. Ben was so thoughtful and thorough in all our conversations, and it really shows in his writing!

Quanta Magazine@quantamagazine.org · 2y ago

Large language models may not be so omnipotent after all. New research shows that LLMs, like humans, prefer to learn some linguistic patterns over others. @benbenbrubaker.bsky.social reports: www.quantamagazine.org/can-ai-model...