Raphaël Millière

@raphaelmilliere.com

Philosopher of Artificial Intelligence & Cognitive Science https://raphaelmilliere.com/

This should be fun! I'll be talking about what it means for a neural network to have a world model (joint work with @raphaelmilliere.com) and presenting new evidence that the standard example, Othello-GPT, does not have one (joint work with @shauli.bsky.social among others).

Eivinas Butkus@eivinas.bsky.social · 2w ago

with Alexa Tartaglini, @zfjoshying.bsky.social, Ellen Su, @solimlegris.bsky.social, @soniajoseph.bsky.social, Patrick Butlin, @johnmorrison.bsky.social, @thisismyhat.bsky.social, @jennhu.bsky.social, Ilker Yildirim, @brendenlake.bsky.social, Niko Kriegeskorte

Sadly won't be at ACL in person, but check out our presentations below! 🌟 I'm giving a (remote) keynote at SCiL on 7/4! 🌟We also have a poster on probability x grammaticality, and a talk on pragmatics x Theory of Mind! Our lab is actively recruiting, so please reach out! Details at glintlab.org

Bild

Now published in open access! Your one-stop shop for the philosophy of language models. It's the spiritual descendant of our two-part preprint from 2024, fully updated. This should be particularly useful for anyone looking for an entry point into this rapidly growing field.

The Philosophy of Language Models

The success of large language models (LLMs) across many domains of AI research has generated intense debate. Some attribute their impressive performance on complex tasks to human-like linguistic and ...

compass.onlinelibrary.wiley.com

Very cool idea! Some quick thoughts. It looks like the corruption preserves a lot of information (function words, morphology, word order, punctuation, numbers, register) which would strongly constrains the posterior over plausible discourse frames as it were. 1/3

Gary Lupyan@glupyan.bsky.social · 7mo ago

arxiv.org/abs/2601.11432 I want to share an astonishing result. LLMs can "translate" Jabberwocky' texts like 'He dwushed a ghanc zawk” & even and even 'In the BLANK BLANK, BLANK BLANK has BLANK over any BLANK BLANK’s BLANK' This has profound consequence for thinking about.. 1/2

Can LLMs reason by analogy like humans? We investigate this question in a new paper published in the Journal of Memory and Language (link below). This was a long-running but very rewarding project. Here are a few thoughts on our methodology and main findings. 1/9

Bild

The sycophantic tone of ChatGPT always sounded familiar, and then I recognized where I'd heard it before: author response letters to reviewer comments. "You're exactly right, that's a great point!" "Thank you so much for this insight!" Also how it always agrees even when it contradicts itself.

Despite extensive safety training, LLMs remain vulnerable to “jailbreaking” through adversarial prompts. Why does this vulnerability persist? In a new paper published in Philosophical Studies, I argue this is because current alignment methods are fundamentally shallow. 🧵 1/13

Bild

Losing Lynch is a strange feeling. This is hardly an original thing to say, but his work left a big impression on me since I was a teenager. I've rewatched most of his movies over the past year, they're every bit as enthralling as I remembered them. Now it's time to rewatch Twin Peaks!

a woman is sitting at a table with the words " one day the sadness will end "

ALT: a woman is sitting at a table with the words " one day the sadness will end "

media.tenor.com

My article 'Constitutive Self-Consciousness' is now published online in the Australasian Journal of Philosophy. It argues (spoiler alert!) against the claim that self-consciousness is constitutive of consciousness.

The claim that consciousness constitutively involves self-consciousness has a long philosophical history, and has received renewed support in recent years. My aim in this paper is to argue that this surprisingly enduring idea is misleading at best, and insufficiently supported at worst. I start by offering an elucidatory account of consciousness, and outlining a number of foundational claims that plausibly follow from it. I subsequently distinguish two notions of self-consciousness: consciousness of oneself and consciousness of one’s experience. While ‘self-consciousness’ is often taken to refer to the former notion, the most common variant of the constitutive claim, on which I focus here, targets the latter. This claim can be further interpreted in two ways: on a deflationary reading, it falls within the scope of foundational claims about consciousness, while on an inflationary reading, it points to determinate aspects of phenomenology that are not acknowledged by the foundational claims as being aspects of all conscious mental states. I argue that the deflationary reading of the constitutive claim is plausible, but should be formulated without using a term as polysemous and suggestive as ‘self-consciousness’; by contrast, the inflationary reading is not adequately supported, and ultimately rests on contentious intuitions about phenomenology. I conclude that we should abandon the idea that self-consciousness is constitutive of consciousness.

I'm happy to share that I'll be one of Schmidt Sciences's new AI2050 fellows! I'll be focusing on addressing the risk of interpretability illusions in AI – cases where interpretability methods yield seemingly plausible yet incorrect explanations. www.schmidtsciences.org/schmidt-scie...

Schmidt Sciences to Award $12 Million to Advance Research on Beneficial AI

AI2050 fellowships recognize scholars working to create AI for a better world

schmidtsciences.org

Three ManyBabies projects - big collaborative replications of infancy phenomena - wrapped up this year. The first paper came out this fall. I thought I'd take this chance to comment on what I make of the non-replication result. 🧵 bsky.app/profile/laur...

Laura Schlingloff-Nemecz@laurasn.bsky.social · 2y ago

The Manybabies4 paper is out! Infants' Social Evaluation of Helpers and Hinderers: A Large-Scale, Multi-Lab, Coordinated Replication Study onlinelibrary.wiley.com/doi/abs/10.1... 1000 babies tested in 37 labs; "Overall, 49.34% of infants preferred Helpers over Hinderers in the social condition"