Max Kleiman-Weiner

@maxkw.bsky.social

professor at university of washington and scientist at Google DeepMind. computational cognitive scientist. working on social and artificial intelligence and alignment. http://faculty.washington.edu/maxkw/

LLMs can shift people's beliefs. But most persuasion studies only check beliefs before and after a conversation. We built PersuasionTrace to measure beliefs turn by turn, so we can study how belief updates actually unfold.

An example human-target persuasion round with multi-turn persuasion tracing.

Task diversity is supposedly key to generalization in RL. But what does it do to continual RL, where agents face one new task distribution after another? We find that past a point, more diversity actually inhibits continual reinforcement learning 🧵

🤔💭What even is reasoning? It's time to answer the hard questions! We built the first unified taxonomy of 28 cognitive elements underlying reasoning Spoiler—LLMs commonly employ sequential reasoning, rarely self-awareness, and often fail to use correct reasoning structures🧠

Bild

New paper challenges how we think about Theory of Mind. What if we model others as executing simple behavioral scripts rather than reasoning about complex mental states? Our algorithm, ROTE (Representing Others' Trajectories as Executables), treats behavior prediction as program synthesis.

Bild
Kunal Jha@kjha02.bsky.social · 10mo ago

Forget modeling every belief and goal! What if we represented people as following simple scripts instead (i.e "cross the crosswalk")? Our new paper shows AI which models others’ minds as Python code 💻 can quickly and accurately predict human behavior! shorturl.at/siUYI%F0%9F%...

When values collide, what do LLMs choose? In our new paper, "Generative Value Conflicts Reveal LLM Priorities," we generate scenarios where values are traded off against each other. We find models prioritize "protective" values in multiple-choice, but shift toward "personal" values when interacting.

Andy Liu@andyliu.bsky.social · 10mo ago

🚨New Paper: LLM developers aim to align models with values like helpfulness or harmlessness. But when these conflict, which values do models choose to support? We introduce ConflictScope, a fully-automated evaluation pipeline that reveals how models rank values under conflict. (📷 xkcd)

Excited by our new work estimating the empowerment of LLM-based agents in text and code. Empowerment is the causal influence an agent has over its environment and measures an agent's capabilities without requiring knowledge of its goals or intentions.

Bild

Claire's new work showing that when an assistant aims to optimize another's empowerment, it can lead to others being disempowered (both as a side effect and as an intentional outcome)!

Claire Yang@claireyang.bsky.social · 12mo ago

Still catching up on my notes after my first #cogsci2025, but I'm so grateful for all the conversations and new friends and connections! I presented my poster "When Empowerment Disempowers" -- if we didn't get the chance to chat or you would like to chat more, please reach out!

Person standing next to poster titled "When Empowerment Disempowers"

Still catching up on my notes after my first #cogsci2025, but I'm so grateful for all the conversations and new friends and connections! I presented my poster "When Empowerment Disempowers" -- if we didn't get the chance to chat or you would like to chat more, please reach out!

Person standing next to poster titled "When Empowerment Disempowers"

Our new paper is out in PNAS: "Evolving general cooperation with a Bayesian theory of mind"! Humans are the ultimate cooperators. We coordinate on a scale and scope no other species (nor AI) can match. What makes this possible? 🧵 www.pnas.org/doi/10.1073/...

Evolving general cooperation with a Bayesian theory of mind | PNAS

Theories of the evolution of cooperation through reciprocity explain how unrelated self-interested individuals can accomplish more together than th...

pnas.org

As always, CogSci has a fantastic lineup of workshops this year. An embarrassment of riches! Still deciding which to pick? If you are interested in building computational models of social cognition, I hope you consider joining @maxkw.bsky.social, @dae.bsky.social, and me for a crash course on memo!

Cognitive Science Society@cogscisociety.bsky.social · last yr.

#Workshop at #CogSci2025 Building computational models of social cognition in memo 🗓️ Wednesday, July 30 📍 Pacifica I - 8:30-10:00 🗣️ Kartik Chandra, Sean Dae Houlihan, and Max Kleiman-Weiner 🧑‍💻 underline.io/events/489/s...

Promotional image for a #CogSci2025 workshop titled “Building computational models of social cognition in memo.” Organized and presented by Kartik Chandra, Sean Dae Houlihan, and Max Kleiman-Weiner. Scheduled for July 30 at 8:30 AM in room Pacifica I. The banner features the conference theme “Theories of the Past / Theories of the Future,” and the dates: July 30–August 2 in San Francisco.

LLMs learn beliefs and values from human data, influence our opinions, and then reabsorb those influenced beliefs, feeding them back to users again and again. We call this the "Lock-In Hypothesis" and develop theory, simulations, and empirics to test it in our latest ICML paper!

Bild

Awesome new work from my lab led by @kjha02.bsky.social scaling cooperative AI! True cooperation requires adapting to both unfamiliar partners and novel environments. Agents trained with CEC get us closer to agents that can act with general cooperative principles rather than memorized strategies.

Kunal Jha@kjha02.bsky.social · last yr.

Our new paper (first one of my PhD!) on cooperative AI reveals a surprising insight: Environment Diversity > Partner Diversity. Agents trained in self-play across many environments learn cooperative norms that transfer to humans on novel tasks. shorturl.at/fqsNN%F0%9F%...