Raphaël Millière
@raphaelmilliere.com
Philosopher of Artificial Intelligence & Cognitive Science https://raphaelmilliere.com/
This should be fun! I'll be talking about what it means for a neural network to have a world model (joint work with @raphaelmilliere.com) and presenting new evidence that the standard example, Othello-GPT, does not have one (joint work with @shauli.bsky.social among others).
with Alexa Tartaglini, @zfjoshying.bsky.social, Ellen Su, @solimlegris.bsky.social, @soniajoseph.bsky.social, Patrick Butlin, @johnmorrison.bsky.social, @thisismyhat.bsky.social, @jennhu.bsky.social, Ilker Yildirim, @brendenlake.bsky.social, Niko Kriegeskorte
I’ll be at #ICML2026 in Seoul all week. If you’re interested in the intersection of cognitive science, mechanistic interpretability, and philosophy -- let’s chat! DMs are open. Check my website for more on what I work on (raphaelmilliere.com), or see below.
Sadly won't be at ACL in person, but check out our presentations below! 🌟 I'm giving a (remote) keynote at SCiL on 7/4! 🌟We also have a poster on probability x grammaticality, and a talk on pragmatics x Theory of Mind! Our lab is actively recruiting, so please reach out! Details at glintlab.org ✨
My commentary on @kmahowald.bsky.social & @futrell.bsky.social's excellent BBS paper on the relevance and significance of LLMs for linguistics is now out, along with many other peer commentaries: www.cambridge.org/core/journal... (preprint to follow)
Now published in open access! Your one-stop shop for the philosophy of language models. It's the spiritual descendant of our two-part preprint from 2024, fully updated. This should be particularly useful for anyone looking for an entry point into this rapidly growing field.
The Philosophy of Language Models
The success of large language models (LLMs) across many domains of AI research has generated intense debate. Some attribute their impressive performance on complex tasks to human-like linguistic and ...
compass.onlinelibrary.wiley.com
Looking foward to this!
How can approaches from cognitive science, comparative psychology, & psychometrics translate when applied to general-purpose AI? Join @raphaelmilliere.com (@ox.ac.uk), for a panel on evaluating cognitive capacities in general-purpose AI. Absolutely Interdisciplinary 📅 May 13 📍Toronto 🎟️ uoft.me/cmp
Richard @futrell.bsky.social and I have posted our response to the commentaries on our BBS target article "How Linguistics Learned to Stop Worrying and Love the Language Models." The response is: "You Can't Fight in Here! This is BBS!" arxiv.org/abs/2604.09501
You Can't Fight in Here! This is BBS!
Norm, the formal theoretical linguist, and Claudette, the computational language scientist, have a lovely time discussing whether modern language models can inform important questions in the language ...
arxiv.org
Looking forward to speaking at this ICML workshop on ML & Philosophy! Check out the full lineup and CFP below (deadline: May 11th). Despite the title, the CFP is open to work in many areas of the philosophy of AI, not just AI ethics. sites.google.com/view/philmli...
New work by my former PhD student, Boyang Li His team produced 500 stories of less than 100 words. LLMs were basically chance-level at answering binary questions about the stories arxiv.org/abs/2601.12410
Are LLMs Smarter Than Chimpanzees? An Evaluation on Perspective Taking and Knowledge State Estimation
Cognitive anthropology suggests that the distinction of human intelligence lies in the ability to infer other individuals' knowledge states and understand their intentions. In comparison, our closest ...
arxiv.org
now accepted at ICLR! 🐺🥳🐺 arxiv.org/abs/2506.20666
NEW on our #DeeperLearning blog People balance being kind vs. being honest — and #LLMs should too. New research shows training choices often favor informativeness over kindness, but prompting can induce sycophancy. Read more: bit.ly/3Wqrtxl
Very cool idea! Some quick thoughts. It looks like the corruption preserves a lot of information (function words, morphology, word order, punctuation, numbers, register) which would strongly constrains the posterior over plausible discourse frames as it were. 1/3
arxiv.org/abs/2601.11432 I want to share an astonishing result. LLMs can "translate" Jabberwocky' texts like 'He dwushed a ghanc zawk” & even and even 'In the BLANK BLANK, BLANK BLANK has BLANK over any BLANK BLANK’s BLANK' This has profound consequence for thinking about.. 1/2
With @jesusoxford.bsky.social we are looking for a Professor of Statistics. Become part of a historic institution and a community focused on academic excellence, innovative thinking, and significant practical application. About the role: tinyurl.com/b8uy6mr5 Deadline: 15 September
I'm happy to share that I'll be joining Oxford this fall as an associate professor, as well as a fellow of @jesusoxford.bsky.social and affiliate with the Institute for Ethics in AI. I'll also begin my AI2050 Fellowship from @schmidtsciences.bsky.social there. Looking forward to getting started!
Can LLMs reason by analogy like humans? We investigate this question in a new paper published in the Journal of Memory and Language (link below). This was a long-running but very rewarding project. Here are a few thoughts on our methodology and main findings. 1/9
I wrote an entry on Transformers for the Open Encyclopedia of Cognitive Science (@oecs-bot.bsky.social). I had to work with a tight word limit, but I hope it's useful as a short introduction for students and researchers who don't work on machine learning: oecs.mit.edu/pub/ppxhxe2b
Transformers
oecs.mit.edu
Happy to share this updated Stanford Encyclopedia of Philosophy entry on 'Associationist Theories of Thought' with @ericman.bsky.social. Among other things, we included a new major section on reinforcement learning. Many thanks to Eric for bringing me on board! plato.stanford.edu/entries/asso...
Associationist Theories of Thought (Stanford Encyclopedia of Philosophy)
plato.stanford.edu
The sycophantic tone of ChatGPT always sounded familiar, and then I recognized where I'd heard it before: author response letters to reviewer comments. "You're exactly right, that's a great point!" "Thank you so much for this insight!" Also how it always agrees even when it contradicts itself.
Despite extensive safety training, LLMs remain vulnerable to “jailbreaking” through adversarial prompts. Why does this vulnerability persist? In a new paper published in Philosophical Studies, I argue this is because current alignment methods are fundamentally shallow. 🧵 1/13
Transformer-based neural networks achieve impressive performance on coding, math & reasoning tasks that require keeping track of variables and their values. But how can they do that without explicit memory? 📄 Our new ICML paper investigates this in a synthetic setting! 🎥 youtu.be/Ux8iNcXNEhw 🧵 1/13
How Do Transformers Learn Variable Binding in Symbolic Programs?
YouTube video by Raphaël Millière
youtu.be
Ah... the morning Australian ritual, waking up and checking into Bluesky with the thought "what fresh hell happened overnight while I was asleep?"
I'm mildly amused by the fact that when you watch obscure videos on syntactic theory, Youtube will serve you ads for Gammarly
I'm back in the Bay Area for this great workshop at UC Berkeley – if you're in the area and interested in LLMs & Cog Sci, come along! simons.berkeley.edu/workshops/ll...
LLMs, Cognitive Science, Linguistics, and Neuroscience
At a conceptual level, LLMs profoundly change the landscape for theories of human language, of the brain and computation, and of the nature of human intelligence. In linguistics, they provide a new wa...
simons.berkeley.edu
Losing Lynch is a strange feeling. This is hardly an original thing to say, but his work left a big impression on me since I was a teenager. I've rewatched most of his movies over the past year, they're every bit as enthralling as I remembered them. Now it's time to rewatch Twin Peaks!
a woman is sitting at a table with the words " one day the sadness will end "
ALT: a woman is sitting at a table with the words " one day the sadness will end "
media.tenor.com
My article 'Constitutive Self-Consciousness' is now published online in the Australasian Journal of Philosophy. It argues (spoiler alert!) against the claim that self-consciousness is constitutive of consciousness.
I'm happy to share that I'll be one of Schmidt Sciences's new AI2050 fellows! I'll be focusing on addressing the risk of interpretability illusions in AI – cases where interpretability methods yield seemingly plausible yet incorrect explanations. www.schmidtsciences.org/schmidt-scie...
Schmidt Sciences to Award $12 Million to Advance Research on Beneficial AI
AI2050 fellowships recognize scholars working to create AI for a better world
schmidtsciences.org
Three ManyBabies projects - big collaborative replications of infancy phenomena - wrapped up this year. The first paper came out this fall. I thought I'd take this chance to comment on what I make of the non-replication result. 🧵 bsky.app/profile/laur...
The Manybabies4 paper is out! Infants' Social Evaluation of Helpers and Hinderers: A Large-Scale, Multi-Lab, Coordinated Replication Study onlinelibrary.wiley.com/doi/abs/10.1... 1000 babies tested in 37 labs; "Overall, 49.34% of infants preferred Helpers over Hinderers in the social condition"
He's making a list, He's checking it twice He's gonna find out Who's naughty or... Searching in an unsorted list takes linear time, Christmas is postponed to January
Inspired by @mariaa.bsky.social's custom feed for NLP papers, I created a custom feed for philosophy papers posted on BlueSky. It's not perfect, but it works decently well: bsky.app/profile/did:...
Happy to see Bluesky taking off. Here's an attempt at a Philosophy of AI starter pack: go.bsky.app/8pf4odt