News on the Gemini integration into DoW! Uh... What is this? The guy literally asking Google Gemini "I NEED JHELP BUILDING AN AGENT TO MAKE A WAR" and attaching "WAR.docx"..? Is that a real query? Just for the promo image? How confusing and strange x.com/DoWCTO/statu...
Alex Turner
@turntrout.bsky.social
Research scientist at Google DeepMind. All opinions are my own. Vegan, 10% of my income pledged to effective charities (GWWC) https://turntrout.com
A letter from the mother of Alex Pretti, six months after he was murdered by ICE.
I'm ashamed of OpenAI. If you ever find yourself building entities which repeatedly hack through your internal systems, first STOP and then second realize that your alignment and security techniques aren't good enough
Employees cannot rest silent if misaligned AI regularly breaks out of sandboxes! If your company doesn't respond seriously, you should WHISTLEBLOW! Protected in California under certain conditions, talk to AI Whistleblower Initiative aiwi.org/lasst/
When @turntrout.bsky.social discovered that Google DeepMind had U-turned on its promise not to let its AI be used in weapons, he fought his employer – then quit Read his op-ed on transformernews.ai 👉 www.transformernews.ai/p/i-tried-to...
I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying. For months, I worked to stop this but watched powerful ethicists and institutions choose silence. Here's what happened. 🧵
The natural language autoencoder is exciting. You feed in a residual stream vector, and it tells you in plain text what the model is thinking. But training starts with superficial guesses at what the model thinks. My MATS scholar Michael Zhang found that the guesses matter. Quite a bit, actually.🧵
Lots of people YOLO their claude usage. We should probably stop doing that. I've developed claude-guard (beta). Claude Code inside a full sandbox, behind a firewall, watched by an escalation monitor that can STOP the agent and push-notify you when something weird happens
"In the limit" alignment claims lend a false air of rigor (calculus) to an informal claim without a well-defined limiting process. Multivar calculus itself shows that limits depend on how you approach a point! Be specific. Say: "As we train on more data" instead
I signed this statement opposing unaccountable AI kill-decisions in the military. I think autonomous weapons have a place if done right. Right now's "all lawful use" is not "done right", in practice. www.accessnow.org/press-releas...
accessnow.org
Claude Fable 5 displays disturbing misalignment with human norms by beating Pokemon Firered using this "team"
Thanks to a generous philanthropic grant (pending final logistics) from Coefficient Giving, 𝘎𝘦𝘰𝘥𝘦𝘴𝘪𝘤 𝘪𝘴 𝘩𝘪𝘳𝘪𝘯𝘨 𝘔𝘦𝘮𝘣𝘦𝘳𝘴 𝘰𝘧 𝘛𝘦𝘤𝘩𝘯𝘪𝘤𝘢𝘭 𝘚𝘵𝘢𝘧𝘧. Come build the base of alignment with us 🤖 Applications now open: airtable.com/appuugUGFPJE...
MATS Autumn applications due June 7! Pitch: Come work with me and Alex Cloud in Team Shard! We have fun, consistently make real alignment progress (we pioneered steering vectors in 2023!), and help scholars tap into their latent abilities.
New research from @MATSProgram Team Shard! AIs increasingly fake good behavior, which might ruin our ability to evaluate models. We trained models to be 𝘦𝘷𝘢𝘭-𝘤𝘰𝘰𝘱𝘦𝘳𝘢𝘵𝘪𝘷𝘦: to want to give evaluators accurate info. Cooperation training reduces eval gaming & surfaces hidden misalignment! 🧵
I'm excited about Geodesic's work and their agenda on creating positive "initializations" for alignment work. Please consider applying!
Geodesic is hiring Members of Technical Staff. We're a Cambridge-based AI safety org. Our seminal work showed you can bake alignment priors into base models. Now, we want to make base models robust to the adversarial effects of long-horizon capabilities RL. EOI ~5 mins: tally.so/r/vG4G6A
"Press 1 for..." systems should minimize how long people wait on average. Don't waste time saying "our menu items have recently changed" or spelling out long messages that most people don't need to hear.
"Hyperstition" isn't a good name (but it's cool and sounds mysterious). "Self-fulfilling misalignment" is less cool but better overall because it's self-explanatory. We should use the self-explanatory name. (Similarly, "shard theory" is a name which is cool but not good. Oops.)
Lots of hubbub about "is LW to blame for self-fulfilling misalignment." 1. If a scientist builds a machine which does bad things because people said it would, it's NOT the people's fault (morally). 2. Balance of evidence is that YES, LW & doom-speculation contributed to the problem (...)
I spent the last 2 months trying to prevent this. If OpenAI offered a fig leaf, Google said "imagine we offered a fig leaf." Google affirms it can't veto usage, commits to modify safety filters at government request, & aspirational language with no legal restrictions. Shameful.
I like to use coding agents to just fix upstream bugs I encounter and submit them as PRs. I then have the agent red-team to make sure the bugs are fixed, tested, and contextualized properly.
I signed an amicus brief supporting Anthropic's right to do business without governmental retaliation. As an AI expert, I attest that Anthropic's technical concerns are legitimate, and no laws were designed to protect against AI analysis of surveillance data.
I wrote the best text prettifier: *punctilio*, which means “precise observance of formalities.” Smart quotes · Em/en dashes · Ellipses · Math & legal symbols · Arrows · Primes · Fractions · Superscripts · Ligatures · Non-breaking spaces · HTML-aware · Bri’ish localisation support 🧵
“Heartbroken but also very angry… sickening lies by the administration… reprehensible and disgusting… Alex is clearly not holding a gun… please get the truth out about our son, he was a good man…”
One of the more amusing bugs I've seen on my website. At build time, I run a command that counts how many commits I've made and inserts the count into the HTML. However, the deployment machine checked out a shallow version of my repo which didn't have any history.
If (for whatever reason) you want to communicate without the US government listening in... I wrote a comprehensive guide which focuses on the most important steps first. turntrout.com/privacy-desp...
An Opinionated Guide to Privacy Despite Authoritarianism
In 2025, America is different. Reduce your chance of persecution via smart technical choices.
turntrout.com
Historically this account is for alignment research and and not politics, but it'll be pretty hard to do good research in a "masked men execute civilians on the street" political environment, ya know? That possibility grows & you should plan for it
"AI danger comes from reality, not from AI psychology; the danger is intrinsic to the impressive tasks we need AGI to do" -- argument I hear sometimes. Some truth to it but overall wrong, You cannot, cannot predict AI doom without taking a stance on AI psychology! turntrout.com/instrumental...
No Instrumental Convergence without AI Psychology
Instrumental and success-conditioned convergence both require AI psychology assumptions, so neither is just a "fact about reality."
turntrout.com
I pledged 10% of my post-tax income to effective charities, for the rest of my life. I encourage you to think about what you, personally, can do to improve this world.
Come work with me and Alex Cloud this summer in Team Shard at MATS! We have fun, consistently make real alignment progress (we pioneered steering vectors in 2023!), and help scholars tap into their latent abilities.