Is it possible to write 100,000 lines of code well, if you do not read it? Let's go Hunting Zombies! davidbau.com/archives/20... In this post I dive into the code of two AI agent contestants in the Teleport coding challenge to learn their secrets. Very fun. And also instructive.
David Bau
@davidbau.bsky.social
Interpretable Deep Networks. http://baulab.info/ @davidbau
Oh man! I love this preprint and also the website Rohit made to demo it. The gaze of a VLM is mediated by a much smaller set of attention heads than the full set, as if "conscious" attention is a small subset of "all attention heads". His demo lets you steer these in realtime.
I recently spoke with Yascha Mounk about how researchers look inside AI to understand how it is thinking. Here is the podcast: writing.yaschamounk.com/p/david-bau-2
David Bau on How—and Whether—Artificial Intelligence Thinks
Yascha Mounk and David Bau examine the mysterious internal processes that drive AI behavior—and why they may be fundamentally alien.
writing.yaschamounk.com
Please join us at NEMI 2026, the 3rd New England Mechanistic Interpretability Workshop! August 14th at Boston University. Register now: nemiconf.github.io/summer26/ A remarkable time for AI. Come share your insights and research on the mechanisms inside our models. bsky.app/profile/mic...
Micah Benson (@micahben.bsky.social)
🧠🤖 The 2026 New England Mechanistic Interpretability (NEMI) Workshop will be Aug. 14 at Boston University! Help spread the word and join the New England mech interp community! Registration and submission info in thread:👇
bsky.app
Musk has no ambition. A 777 has 2.6 million lines of code *not* because that's what it takes to fly. It's because that's what it takes to lift 370 tons safely from LAX to Heathrow every day without endangering its 400 passengers. www.youtube.com/watch?v=SZk...
Coding careers will die by Dec 2026: Elon Musk
Traditional coding as a job is standing on a burning platform. In...
youtube.com
"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.
The Teleport Contest is open. Port NetHack 5.0 from C to JavaScript, bit-exactly. Same screen, every keystroke. Any approach: LLM agents, hand-coded, transpiler, hybrid. Live leaderboard, two phases through December. mazesofmenace.ai/announcement
NetHack is one of the most complex and longest-lived open source programs ever written, and after 46 years, v5.0 shipped today. www.nethack.org/common/inde... And ... it is a VERY cool large codebase to work with in the LLM era.
2026 is a whirlwind year for AI. Underlying it all is the greatest scientific mystery of our age. How does a neural network think? I talked w Oliver Whang in NYTimes Magazine, on how AI interpretability is a tangle of structure waiting to be unraveled: www.nytimes.com/2026/04/15/...
Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top".
Calling attention to an exciting "deception detection" hackathon we're planning this summer! w @NDIF and @CadenzaLabs. Recruiting red teams now, blue teams later. Red teams, time is short: proposals due Mar 31. $10K stipend + compute, $15K finals prize. nnsight.net/blog/2026/0...
In 1982, high school students in Sudbury, Mass. wrote a dungeon game called Hack. They had Atari 800s and Logo and an obsession with a Unix game called Rogue that most of them had never seen. I grew up one town over with the same computers and the same obsession.
Sam Altman and Dario Amodei have both staked out positions on AI weapons. But you can see from what they've said: the gap between them is a question of professional ethics. bsky.app/profile/mas...
Mike Masnick (@masnick.com)
My goodness.
bsky.app
I will be adding some time in my AI research group today for researchers and engineers to discuss the mission and ethics of all our work. We are often too preoccupied by the details. Good work requires clear purpose. Today is a good day to reflect.
Those of us who work in AI in the US today should take a moment to think today. Do not get distracted by the circus. Instead, let us pause to think carefully about our freedoms, our rights, and our responsibilities as citizens and professionals. It is a deadly serious moment.
Those of us who work in AI in the US today should take a moment to think today. Do not get distracted by the circus. Instead, let us pause to think carefully about our freedoms, our rights, and our responsibilities as citizens and professionals. It is a deadly serious moment.
Themes
Are we all Agents of Chaos in AI? (Hope not!) In recent weeks using OpenClaw has taught us a lot about this wooly new kind of autonomous software agent. Its valuable to see what @NatalieShapira, @wendlerch et al. have seen: agentsofchaos.baulab.info/
How do you knock the induction heads out of an LM while preserving its ability to think? Is it even possible? @keremsahin22.bsky.social's work is worth reading if you haven't seen it yet. hapax.baulab.info
The Art of Wanting. About the question I see as central in AI ethics, interpretability, and safety. Can an AI take responsibility? I do not think so, but *not* because it's not smart enough. davidbau.com/archives/20...
I think everyone (not just academics) should read this.
What should academics be doing right now? I have been writing up some thoughts on what the research says about effective action, and what universities specifically can do. davidbau.github.io/poetsandnurs... It's on GitHub. Suggestions and pull requests welcome. github.com/davidbau/poe...
What should academics be doing right now? I have been writing up some thoughts on what the research says about effective action, and what universities specifically can do. davidbau.github.io/poetsandnurs... It's on GitHub. Suggestions and pull requests welcome. github.com/davidbau/poe...
From induction to FVs, every ICL mechanism we've pinned down is fuzzy copying. Is copying all there is? @ericwtodd.bsky.social trained on groups where tokens have no fixed meaning and found a basket of mechanisms beyond copying. Watch them emerge, a grokking cascade! ↓ bsky.app/profile/eri...
I can't read Chinese, but my family has old genealogy documents I've always wanted to understand. Claude and Gemini helped me build an interactive reader to explore the calligraphy character by character. I can finally read my great-grandfather's epitaph. Try it: davidbau.com/archives/202...
My vibe-coded Mandelbrot viewer is 40x faster now! New GPU synchronization tricks go outside the design intent of WebGPU specs. But the real story: Claude tells me what happens in the AGI break room. What superhuman AGIs say when the boss is not around: davidbau.com/archives/202...
I have been teaching myself to vibe code. Watch Claude Code grow my 780 lines to 13,600 - mandelbrot.page/coverage/ca... Two fundamental rules for staying in control: davidbau.com/archives/20...
I have been teaching myself to vibe code. Watch Claude Code grow my 780 lines to 13,600 - mandelbrot.page/coverage/ca... Two fundamental rules for staying in control: davidbau.com/archives/20...
At the #Neurips2025 mechanistic interpretability workshop I gave a brief talk about Venetian glassmaking, since I think we face a similar moment in AI research today. Here is a blog post summarizing the talk: davidbau.com/archives/202...
The secret life of an LM is defined by its internal data types. Inner layers transport abstractions that are more robust than words, like concepts, functions, or pointers. In new work yesterday, @arnabsensharma.bsky.social et al identify a data type for *predicates*. bsky.app/profile/arn...
Arnab Sen Sharma (@arnabsensharma.bsky.social)
How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵
bsky.app
What does an LLM do when it translates from Italian "amore" to Spanish "amor" or French "amour"? That's easy! (you might think) Because surely it knows: amore, amor, amour are all based on the same Latin word. It can just drop the "e", or add a "u".
Looking forward to #COLM2025 tomorrow. DM me if you'll also be there and want to meet to chat.
Who is going to be at #COLM2025? I want to draw your attention to a COLM paper by my student @sfeucht.bsky.social that has totally changed the way I think and teach about LLM representations. The work is worth knowing. And you can meet Sheridan at COLM, Oct 7! bsky.app/profile/sfe...
There are a lot of interesting details that surface when you use SAEs to understand and control diffusion image synthesis models. Learn more in @wendlerc.bsky.social's talk.
New YouTube video posted! @wendlerc.bsky.social presents his work using SAEs for diffusion text-to-image models. The authors find interpretable SAE features and demonstrate how these features can alter generated images. Watch here: youtu.be/43NnaqGjArA
On the Good Fight podcast w substack.com/@yaschamounk I give a quick but careful primer on how modern AI works. I also chat about our responsibility as machine learning scientists, and what we need to fix to get AI right. Take a listen and reshare - www.persuasion.community/p/david-bau
David Bau on How Artificial Intelligence Works
Yascha Mounk and David Bau delve into the “black box” of AI.
persuasion.community