Just arrived in Rio for #CogSci2026! I'll be at the Cognitive Science of AI Alignment workshop on Wednesay afternoon to talk about "Machines That Care Like Us"
Max Kleiman-Weiner
@maxkw.bsky.social
professor at university of washington and scientist at Google DeepMind. computational cognitive scientist. working on social and artificial intelligence and alignment. http://faculty.washington.edu/maxkw/
Excited about our new work measuring multi-turn persuasion in AI-human interactions and how to simulate human persuadability!
LLMs can shift people's beliefs. But most persuasion studies only check beliefs before and after a conversation. We built PersuasionTrace to measure beliefs turn by turn, so we can study how belief updates actually unfold.
LLMs can shift people's beliefs. But most persuasion studies only check beliefs before and after a conversation. We built PersuasionTrace to measure beliefs turn by turn, so we can study how belief updates actually unfold.
Task diversity is supposedly key to generalization in RL. But what does it do to continual RL, where agents face one new task distribution after another? We find that past a point, more diversity actually inhibits continual reinforcement learning 🧵
Really excited to have the opportunity to give a talk on this work @cogscisociety.bsky.social !!! last year was a blast can’t wait to go back to Rio in July 🇧🇷 HUGE thanks to my collaborators for the support @aydanhuang265.bsky.social @EricYe29011995 @natashajaques.bsky.social @maxkw.bsky.social 🙏
Forget modeling every belief and goal! What if we represented people as following simple scripts instead (i.e "cross the crosswalk")? Our new paper shows AI which models others’ minds as Python code 💻 can quickly and accurately predict human behavior! shorturl.at/siUYI%F0%9F%...
Our new short piece in TiCS on intuitive theories of truth: how people judge whether statements could be true, whether statements are true, and whether to assert them as true. A great collab with @keremoktar.bsky.social @ihandleyminer.bsky.social @kevinzollman.com @lianeleeyoung.bsky.social
New paper: Intuitive theories of truth We connect philosophical theories of truth with cognitive science. We suggest new avenues for research around questions of how people judge statements as truth apt, what makes them true, and whether to assert something as true. Check it out!
Can't wait to present this work @iclr-conf.bsky.social this year!!! Looking forward to hearing everyone's thoughts on the paper and learning more about peoples' research! Thanks again to my collaborators for all of their help on this project!
Forget modeling every belief and goal! What if we represented people as following simple scripts instead (i.e "cross the crosswalk")? Our new paper shows AI which models others’ minds as Python code 💻 can quickly and accurately predict human behavior! shorturl.at/siUYI%F0%9F%...
🤔💭What even is reasoning? It's time to answer the hard questions! We built the first unified taxonomy of 28 cognitive elements underlying reasoning Spoiler—LLMs commonly employ sequential reasoning, rarely self-awareness, and often fail to use correct reasoning structures🧠
Forget modeling every belief and goal! What if we represented people as following simple scripts instead (i.e "cross the crosswalk")? Our new paper shows AI which models others’ minds as Python code 💻 can quickly and accurately predict human behavior! shorturl.at/siUYI%F0%9F%...
New paper challenges how we think about Theory of Mind. What if we model others as executing simple behavioral scripts rather than reasoning about complex mental states? Our algorithm, ROTE (Representing Others' Trajectories as Executables), treats behavior prediction as program synthesis.
Forget modeling every belief and goal! What if we represented people as following simple scripts instead (i.e "cross the crosswalk")? Our new paper shows AI which models others’ minds as Python code 💻 can quickly and accurately predict human behavior! shorturl.at/siUYI%F0%9F%...
When values collide, what do LLMs choose? In our new paper, "Generative Value Conflicts Reveal LLM Priorities," we generate scenarios where values are traded off against each other. We find models prioritize "protective" values in multiple-choice, but shift toward "personal" values when interacting.
🚨New Paper: LLM developers aim to align models with values like helpfulness or harmlessness. But when these conflict, which values do models choose to support? We introduce ConflictScope, a fully-automated evaluation pipeline that reveals how models rank values under conflict. (📷 xkcd)
Excited by our new work estimating the empowerment of LLM-based agents in text and code. Empowerment is the causal influence an agent has over its environment and measures an agent's capabilities without requiring knowledge of its goals or intentions.
Claire's new work showing that when an assistant aims to optimize another's empowerment, it can lead to others being disempowered (both as a side effect and as an intentional outcome)!
Still catching up on my notes after my first #cogsci2025, but I'm so grateful for all the conversations and new friends and connections! I presented my poster "When Empowerment Disempowers" -- if we didn't get the chance to chat or you would like to chat more, please reach out!
Still catching up on my notes after my first #cogsci2025, but I'm so grateful for all the conversations and new friends and connections! I presented my poster "When Empowerment Disempowers" -- if we didn't get the chance to chat or you would like to chat more, please reach out!
lol this may be the most cogsci cogsci slide I've ever seen, from @maxkw.bsky.social "before I got married I had six theories about raising children, now I have six kids and no theories"......but here's another theory #cogsci2025
Our new paper is out in PNAS: "Evolving general cooperation with a Bayesian theory of mind"! Humans are the ultimate cooperators. We coordinate on a scale and scope no other species (nor AI) can match. What makes this possible? 🧵 www.pnas.org/doi/10.1073/...
Evolving general cooperation with a Bayesian theory of mind | PNAS
Theories of the evolution of cooperation through reciprocity explain how unrelated self-interested individuals can accomplish more together than th...
pnas.org
As always, CogSci has a fantastic lineup of workshops this year. An embarrassment of riches! Still deciding which to pick? If you are interested in building computational models of social cognition, I hope you consider joining @maxkw.bsky.social, @dae.bsky.social, and me for a crash course on memo!
#Workshop at #CogSci2025 Building computational models of social cognition in memo 🗓️ Wednesday, July 30 📍 Pacifica I - 8:30-10:00 🗣️ Kartik Chandra, Sean Dae Houlihan, and Max Kleiman-Weiner 🧑💻 underline.io/events/489/s...
Very excited for this workshop!
#Workshop at #CogSci2025 Building computational models of social cognition in memo 🗓️ Wednesday, July 30 📍 Pacifica I - 8:30-10:00 🗣️ Kartik Chandra, Sean Dae Houlihan, and Max Kleiman-Weiner 🧑💻 underline.io/events/489/s...
#Workshop at #CogSci2025 Building computational models of social cognition in memo 🗓️ Wednesday, July 30 📍 Pacifica I - 8:30-10:00 🗣️ Kartik Chandra, Sean Dae Houlihan, and Max Kleiman-Weiner 🧑💻 underline.io/events/489/s...
'Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination' @kjha02.bsky.social · Wilka Carvalho · Yancheng Liang · Simon Du · @maxkw.bsky.social · @natashajaques.bsky.social doi.org/10.48550/arX... (3/20)
Settling in for my flight and apparently A.I. DOOM is now a movie genre between Harry Potter and Classics. Nothing better than an existential crisis with pretzels and a ginger ale.
Thanks to the Diverse Intelligence Community for all these inspiring days & impressions in Sydney 🙏🏻 @chriskrupenye.bsky.social @katelaskowski.bsky.social @divintelligence.bsky.social @maxkw.bsky.social
LLMs learn beliefs and values from human data, influence our opinions, and then reabsorb those influenced beliefs, feeding them back to users again and again. We call this the "Lock-In Hypothesis" and develop theory, simulations, and empirics to test it in our latest ICML paper!
Excited to speak about some new work on Bayesian Cooperation at this workshop! Join us virtually
I am really happy to share information about the 2nd Workshop on Modeling and Applications of Evolutionary Game Theory, which will be held virtually on Thursday, May 8, and Friday, May 9. sites.google.com/view/2nd-evo... I enjoyed organizing this workshop with Olivia Chu and Alex McAvoy.
Now out in JPSP ‼️ "Inference from social evaluation" with Zach Davis, Kelsey Allen, @maxkw.bsky.social, and @julianje.bsky.social 📃 (paper): psycnet.apa.org/record/2026-... 📜 (preprint): osf.io/preprints/ps...
Our new paper (first one of my PhD!) on cooperative AI reveals a surprising insight: Environment Diversity > Partner Diversity. Agents trained in self-play across many environments learn cooperative norms that transfer to humans on novel tasks. shorturl.at/fqsNN%F0%9F%...
Awesome new work from my lab led by @kjha02.bsky.social scaling cooperative AI! True cooperation requires adapting to both unfamiliar partners and novel environments. Agents trained with CEC get us closer to agents that can act with general cooperative principles rather than memorized strategies.
Our new paper (first one of my PhD!) on cooperative AI reveals a surprising insight: Environment Diversity > Partner Diversity. Agents trained in self-play across many environments learn cooperative norms that transfer to humans on novel tasks. shorturl.at/fqsNN%F0%9F%...
How AlphaGo like architectures can explain human insight. Out now in Cognition!
my paper with max, @maxkw.bsky.social, tuomas, and @fierycushman.bsky.social out in cognition at long last www.sciencedirect.com/science/arti... We explain why humans and successful AI planners both fail on a certain kind of problem that we might describe as requiring insight or creativity
my paper with max, @maxkw.bsky.social, tuomas, and @fierycushman.bsky.social out in cognition at long last www.sciencedirect.com/science/arti... We explain why humans and successful AI planners both fail on a certain kind of problem that we might describe as requiring insight or creativity
Similar failures of consideration arise in human and machine planning
Humans are remarkably efficient at decision making, even in “open-ended” problems where the set of possible actions is too large for exhaustive evalua…
sciencedirect.com