Human thought is thought to be systematic: Do reasoning models have systematicity of thought? With @brendenlake.bsky.social we study this question in our new preprint arxiv.org/abs/2609.13948 Thread 🧵
Siyuan Song
@siyuansong.bsky.social
Grad student@Princeton Psychology siyuansong.site Language, Learning, Intelligence Prev: Undergrad@UTexas, SJTU Summer Research Visit @MIT BCS, Harvard Psych Opinions are my own.
For a year and a half, @carorowland.bsky.social, @lehersingh.bsky.social, Marisa Casillas, Shanley Allen, and I have been meeting to discuss whether innateness is still a useful concept to think about in studying language acquisition. Here's our take: osf.io/preprints/ps...
Chinese babyLM is on folks:
🚀 Announcing the Chinese BabyLM Challenge: the first shared task on data-efficient pretraining for Chinese. 📍 Co-located with NLPCC 2026 (Nov 3–5, Macau🇨🇳🇲🇴) Can you train a strong Chinese LM on just ~100M words? chinese-babylm.github.io 🧵 👇(1/6)
🚀 Announcing the Chinese BabyLM Challenge: the first shared task on data-efficient pretraining for Chinese. 📍 Co-located with NLPCC 2026 (Nov 3–5, Macau🇨🇳🇲🇴) Can you train a strong Chinese LM on just ~100M words? chinese-babylm.github.io 🧵 👇(1/6)
Chinese BabyLM Challenge
chinese-babylm.github.io
Just arrived in Boston for #HSP2026! I'll be presenting my work with @thomashikaru.bsky.social on error sensitivity in next-word predictions of humans and LMs — Friday 12:10–2:00pm poster session. Come say hi!
I'm hiring a new lab manager for my lab @ UCSD! For more info on the lab, check out our website: lillab.ucsd.edu Target start date is June 1 (flexible) and application deadline is March 26. Please share with anyone you think might be a good fit! Apply here: employment.ucsd.edu/laboratory-c...
Laboratory Coordinator - 138788
Laboratory Coordinator - 138788 | Careers at UC San Diego
employment.ucsd.edu
What is the interplay between representations learned from (language) surface forms alone, and those learned from more grounded evidence (e.g.,vision)? Excited to share new work understanding “Cross-modal taxonomic generalization” in (V)LMs arxiv.org/abs/2603.07474 1/
Can large language models *introspect*? In a new paper, @kmahowald.bsky.social and I study the MECHANISM of introspection in big open-source models. tldr: Models detect internal anomalies through DIRECT ACCESS, but don't know what the anomalies are. And they love to guess “apple” 🍎
“All bears have a property”, “Some bears have a property”, “Bears have a property” are different in terms of how the property is generalized to a specific bear – a great example of how language constrains thought! This holds for kids, adults, and according to our new work, (V)LMs! 🧵
Our first South by Semantics lecture of the semester at UT Austin is happening next week on January 30th! I'm excited to hear Dr. Amir Zeldes (Associate Professor at Georgetown University) talk about saliency in discourse and the memorability of salient information for both humans and LLMs.
🧑🔬I’m recruiting PhD students in Natural Language Processing @unileipzig.bsky.social Computer Science, together with @scadsai.bsky.social! Topics include, but aren’t limited to: 🔎Linguistic Interpretability 🌍Multilingual Evaluation 📖Computational Typology Please share! #NLProc #NLP
Can we use VLMs to quantify multimodal alignment in children's experiences? We analyze a large corpus of headcam videos to find out! New preprint from our BabyView project, led by @alvinwmtan.bsky.social and Jane Yang: arxiv.org/abs/2511.18824
Looking forward to #NeurIPS25 this week 🏝️! I'll be presenting at Poster Session 3 (11-2 on Thursday). Feel free to reach out!
Excited to announce that I’ll be presenting a paper at #NeurIPS this year! Reach out if you’re interested in chatting about LM training dynamics, architectural differences, shortcuts/heuristics, or anything at the CogSci/NLP/AI interface in general! #Neurips2025
I’m excited to present SimpleStories at EurIPS! Also if anyone at #EurIPS is interested in chatting about LLM data efficiency, interpretability, model inconsistency or other topics feel free to DM me. Dataset and models: lnkd.in/e_VGWqhP Code: lnkd.in/eEidmv74 Paper: lnkd.in/eH6jS9uY
String probability might be the best tool for assessing LMs' grammatical knowledge, yet it does not directly tell you 'how grammatical' a string is. Here's why and how we should use string probability and minimal pairs: Excited to see this out - it's my great honor to be part of this amazing team!
New work to appear @ TACL! Language models (LMs) are remarkably good at generating novel well-formed sentences, leading to claims that they have mastered grammar. Yet they often assign higher probability to ungrammatical strings than to grammatical strings. How can both things be true? 🧵👇
Oh cool! Excited this LM + construction paper was SAC-Highlighted! Check it out to see how LM-derived measures of statistical affinity separate out constructions with similar words like "I was so happy I saw you" vs "It was so big it fell over".
Josh Rozner's paper (w/ @rifter.bsky.social + @kmahowald.bsky.social) was an SAC Highlight at #EMNLP25! aclanthology.org/2025.emnlp-m...
Delighted Sasha's (first year PhD!) work using mech interp to study complex syntax constructions won an Outstanding Paper Award at EMNLP! Also delighted the ACL community continues to recognize unabashedly linguistic topics like filler-gaps... and the huge potential for LMs to inform such topics!
aclanthology.org
A key hypothesis in the history of linguistics is that different constructions share underlying structure. We take advantage of recent advances in mechanistic interpretability to test this hypothesis in Language Models. New work with @kmahowald.bsky.social and @cgpotts.bsky.social! 🧵👇!
Interested in doing a PhD at the intersection of human and machine cognition? ✨ I'm recruiting students for Fall 2026! ✨ Topics of interest include pragmatics, metacognition, reasoning, & interpretability (in humans and AI). Check out JHU's mentoring program (due 11/15) for help with your SoP 👇
The department of Cognitive Science @jhu.edu is seeking motivated students interested in joining our interdisciplinary PhD program! Applications due 1 Dec Our PhD students also run an application mentoring program for prospective students. Mentoring requests due November 15. tinyurl.com/2nrn4jf9
🧠 New at #NeurIPS2025! 🎵 We're far from the shallow now🎵 TL;DR: We introduce the first "reasoning embedding" and uncover its unique spatio-temporal pattern in the brain. 🔗 arxiv.org/abs/2510.228...
Introducing Global PIQA, a new multilingual benchmark for 100+ languages. This benchmark is the outcome of this year’s MRL shared task, in collaboration with 300+ researchers from 65 countries. This dataset evaluates physical commonsense reasoning in culturally relevant contexts.
Very excited to be going to Chicago for @agnescallard.bsky.social's famous Night Owls next week! I'll be discussing my essay "ChatGPT and the Meaning of Life". Hope to see you there if you're local!
If I spill the tea—“Did you know Sue, Max’s gf, was a tennis champ?”—but then if you reply “They’re dating?!” I’d be a bit puzzled, since that’s not the main point! Humans can track what’s ‘at issue’ in conversation. How sensitive are LMs to this distinction? New paper w/ @sangheekim.bsky.social!
I will be recruiting PhD students via Georgetown Linguistics this application cycle! Come join us in the PICoL (pronounced “pickle”) lab. We focus on psycholinguistics and cognitive modeling using LLMs. See the linked flyer for more details: bit.ly/3L3vcyA
"Although I hate leafy vegetables, I prefer daxes to blickets." Can you tell if daxes are leafy vegetables? LM's can't seem to! 📷 We investigate if LMs capture these inferences from connectives when they cannot rely on world knowledge. New paper w/ Daniel, Will, @jessyjli.bsky.social
Honored to get the chance to contribute to the Chinese dataset! And had a great time working with all the awesome collaborators!
🌍Introducing BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data! LLMs learn from vastly more data than humans ever experience. BabyLM challenges this paradigm by focusing on developmentally plausible data We extend this effort to 45 new languages!
Excited to present this at COLM tomorrow! (Tuesday, 11:00 AM poster session)
One of the ways that LLMs can be inconsistent is the "generator-validator gap," where LLMs deem their own answers incorrect. 🎯 We demonstrate that ranking-based discriminator training can significantly reduce this gap, and improvements on one task often generalize to others! 🧵👇
I will be giving a short talk on this work at the COLM Interplay workshop on Friday (also to appear at EMNLP)! Will be in Montreal all week and excited to chat about LM interpretability + its interaction with human cognition and ling theory.
A key hypothesis in the history of linguistics is that different constructions share underlying structure. We take advantage of recent advances in mechanistic interpretability to test this hypothesis in Language Models. New work with @kmahowald.bsky.social and @cgpotts.bsky.social! 🧵👇!
Traveling to my first @colmweb.org🍁 Not presenting anything but here are two posters you should visit: 1. @qyao.bsky.social on Controlled rearing for direct and indirect evidence for datives (w/ me, @weissweiler.bsky.social and @kmahowald.bsky.social), W morning Paper: arxiv.org/abs/2503.20850
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
Language models (LMs) tend to show human-like preferences on a number of syntactic phenomena, but the extent to which these are attributable to direct exposure to the phenomena or more general propert...
arxiv.org
On my way to #COLM2025 🍁 Check out jessyli.com/colm2025 QUDsim: Discourse templates in LLM stories arxiv.org/abs/2504.09373 EvalAgent: retrieval-based eval targeting implicit criteria arxiv.org/abs/2504.15219 RoboInstruct: code generation for robotics with simulators arxiv.org/abs/2405.20179
I’m at #COLM2025 from Wed with: @siyuansong.bsky.social Tue am introspection arxiv.org/abs/2503.07513 @qyao.bsky.social Wed am controlled rearing: arxiv.org/abs/2503.20850 @sashaboguraev.bsky.social INTERPLAY ling interp: arxiv.org/abs/2505.16002 I’ll talk at INTERPLAY too. Come say hi!
Language Models Fail to Introspect About Their Knowledge of Language
There has been recent interest in whether large language models (LLMs) can introspect about their own internal states. Such abilities would make LLMs more interpretable, and also validate the use of s...
arxiv.org