Excited about new work from my CCN keynote (38m): youtu.be/v3J-vJfxhOE?... We often choose between Bayesian and neural net models of behavior. But what if there's a spectrum, and combining both makes better predictions? Introducing BBT w @akjagadish.bsky.social & Guangyuan arxiv.org/abs/2608.22154
Brenden Lake
@brendenlake.bsky.social
Associate Professor of Computer Science and Psychology @ Princeton. Posts are my views only. https://www.cs.princeton.edu/~bl8144/
Day 2 of #CCN2026. @brendenlake.bsky.social kicks us off with a Keynote where he puts the drama of AI companies aside to ask what can AI do for us (cognitive computational scientists)
Check out @smonsays.bsky.social's work on detecting LLMs in experiments based on their lack of human memory constraints
LLM agents are a serious problem for online experiments. It is very easy to use them and very hard to spot them. What can researchers do? With @brendenlake.bsky.social, we suggest detecting LLMs based on their lack of human cognitive constraints in our #CogSci2026 paper: arxiv.org/abs/2604.00016
The best write up on the state of AI in a while, from James Somers, with input from my Princeton colleagues Ken Norman, Uri Hasson, Jon Cohen, and many others. Coding with LLMs was a striking moment for me too www.newyorker.com/magazine/202...
The Case That A.I. Is Thinking
ChatGPT does not have an inner life. Yet it seems to know what it’s talking about.
newyorker.com
There are still open desks in our new Human & Machine Intelligence lab at Princeton. Express your interest in joining us: lake-lab.github.io/apply/
Today in Nature Machine Intelligence, Kazuki Irie & I discuss 4 classic challenges for neural nets — systematic generalization, catastrophic forgetting, few-shot learning, & reasoning. We argue there is a unifying fix: the right incentives & practice. rdcu.be/eLRmg
Our new lab for Human & Machine Intelligence is officially open at Princeton University! Consider applying for a PhD or Postdoc position, either through Computer Science or Psychology. You can register interest on our new website lake-lab.github.io (1/2)
For much more, see the paper! arxiv.org/abs/2508.05776 By Tom Griffiths, Brenden Lake, Tom McCoy, Ellie Pavlick, and Taylor Webb (@cocoscilab.bsky.social, @brendenlake.bsky.social, @rtommccoy.bsky.social, Ellie Pavlick, @taylorwwebb.bsky.social) 9/9
Whither symbols in the era of advanced neural networks?
Some of the strongest evidence that human minds should be thought about in terms of symbolic systems has been the way they combine ideas, produce novelty, and learn quickly. We argue that modern neura...
arxiv.org
I'm joining Princeton University as an Associate Professor of Computer Science and Psychology this fall! Princeton is ambitiously investing in AI and Natural & Artificial Minds, and I'm excited for my lab to contribute. Recruiting postdocs and Ph.D. students in CS and Psychology — join us!
Fantastic new work by @johnchen6.bsky.social (with @brendenlake.bsky.social and me trying not to cause too much trouble). We study systematic generalization in a safety setting and find LLMs struggle to consistently respond safely when we vary how we ask naive questions. More analyses in the paper!
Do LLMs show systematic generalization of safety facts to novel scenarios? Introducing our work SAGE-Eval, a benchmark consisting of 100+ safety facts and 10k+ scenarios to test this! - Claude-3.7-Sonnet passes only 57% of facts evaluated - o1 and o3-mini passed <45%! 🧵
Failures of systematic generalization in LLMs can lead to real-world safety issues. New paper by @johnchen6.bsky.social and @guydav.bsky.social, arxiv.org/abs/2505.21828
Do LLMs show systematic generalization of safety facts to novel scenarios? Introducing our work SAGE-Eval, a benchmark consisting of 100+ safety facts and 10k+ scenarios to test this! - Claude-3.7-Sonnet passes only 57% of facts evaluated - o1 and o3-mini passed <45%! 🧵
Before LLMs, neural nets were task-specific (while humans were task-general). Shockingly, LLMs changed that. How do LLMs represent a task, and do different prompts lead to the same task rep.? Love this by @guydav.bsky.social, and the function vectors of @ericwtodd.bsky.social @davidbau.bsky.social
New preprint alert! We often prompt ICL tasks using either demonstrations or instructions. How much does the form of the prompt matter to the task representation formed by a language model? Stick around to find out 1/N
🤔 Interested in models of social interaction and computational psychiatry? 🤗 If so, @shawnrhoadsphd.bsky.social and I are seeking a highly motivated and talented postdoc to work on these topics! Please share widely! apply.interfolio.com/165809
Despite the world being on fire, I can't help but be thrilled to announce that I'll be starting as an Assistant Professor in the Cognitive Science Program at Dartmouth in Fall '26. I'll be recruiting grad students this upcoming cycle—get in touch if you're interested!
New work in Nature Machine Intelligence by @guydav.bsky.social, @brendenlake.bsky.social, Todd Gureckis, Graham Todd, and Julian Togelius models how humans develop goals—research that could help bridge the gap between human intentions and AI systems. nyudatascience.medium.com/what-is-a-go...
What is a Goal? Advancing Machine Agency While Understanding Human Goal Creation
New CDS research models how people understand and formulate goals — knowledge that could improve AI alignment
nyudatascience.medium.com
I snuck a moment with my son Logan (2.5), ever the creative goal generator, into Fig. 1: "Papa, I made a Truck Carrier Truck!" How do people compose existing concepts to create new goals? Can models generate and understand goals too? nature.com/articles/s4225
Out today in Nature Machine Intelligence! From childhood on, people can create novel, playful, and creative goals. Models have yet to capture this ability. We propose a new way to represent goals and report a model that can generate human-like goals in a playful setting... 1/N
@solimlegris.bsky.social and Wai Keen Vong estimated that average human performance on ARC is about 64%(public eval set). Thus, o3 is clearly better than the average crowd worker tested. Note that almost all tasks were solvable by at least one person who tried it on MTurk. arxiv.org/abs/2409.01374
I'm new here. I heard bluesky is like science Twitter back in the day, and there are fewer posts from Elon Musk. Did I come to the right place?