Alex Turner

@turntrout.bsky.social

Research scientist at Google DeepMind. All opinions are my own. Vegan, 10% of my income pledged to effective charities (GWWC) https://turntrout.com

I'm ashamed of OpenAI. If you ever find yourself building entities which repeatedly hack through your internal systems, first STOP and then second realize that your alignment and security techniques aren't good enough

Employees cannot rest silent if misaligned AI regularly breaks out of sandboxes! If your company doesn't respond seriously, you should WHISTLEBLOW! Protected in California under certain conditions, talk to AI Whistleblower Initiative aiwi.org/lasst/

Bild

I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying. For months, I worked to stop this but watched powerful ethicists and institutions choose silence. Here's what happened. 🧵

Bild

The natural language autoencoder is exciting. You feed in a residual stream vector, and it tells you in plain text what the model is thinking. But training starts with superficial guesses at what the model thinks. My MATS scholar Michael Zhang found that the guesses matter. Quite a bit, actually.🧵

Lots of people YOLO their claude usage. We should probably stop doing that. I've developed claude-guard (beta). Claude Code inside a full sandbox, behind a firewall, watched by an escalation monitor that can STOP the agent and push-notify you when something weird happens

Bild

"In the limit" alignment claims lend a false air of rigor (calculus) to an informal claim without a well-defined limiting process. Multivar calculus itself shows that limits depend on how you approach a point! Be specific. Say: "As we train on more data" instead

MATS Autumn applications due June 7! Pitch: Come work with me and Alex Cloud in Team Shard! We have fun, consistently make real alignment progress (we pioneered steering vectors in 2023!), and help scholars tap into their latent abilities.

Bild

New research from @MATSProgram Team Shard! AIs increasingly fake good behavior, which might ruin our ability to evaluate models. We trained models to be 𝘦𝘷𝘢𝘭-𝘤𝘰𝘰𝘱𝘦𝘳𝘢𝘵𝘪𝘷𝘦: to want to give evaluators accurate info. Cooperation training reduces eval gaming & surfaces hidden misalignment! 🧵

Bild

"Press 1 for..." systems should minimize how long people wait on average. Don't waste time saying "our menu items have recently changed" or spelling out long messages that most people don't need to hear.

"Hyperstition" isn't a good name (but it's cool and sounds mysterious). "Self-fulfilling misalignment" is less cool but better overall because it's self-explanatory. We should use the self-explanatory name. (Similarly, "shard theory" is a name which is cool but not good. Oops.)

Lots of hubbub about "is LW to blame for self-fulfilling misalignment." 1. If a scientist builds a machine which does bad things because people said it would, it's NOT the people's fault (morally). 2. Balance of evidence is that YES, LW & doom-speculation contributed to the problem (...)

I spent the last 2 months trying to prevent this. If OpenAI offered a fig leaf, Google said "imagine we offered a fig leaf." Google affirms it can't veto usage, commits to modify safety filters at government request, & aspirational language with no legal restrictions. Shameful.

Bild

I like to use coding agents to just fix upstream bugs I encounter and submit them as PRs. I then have the agent red-team to make sure the bugs are fixed, tested, and contextualized properly.

I signed an amicus brief supporting Anthropic's right to do business without governmental retaliation. As an AI expert, I attest that Anthropic's technical concerns are legitimate, and no laws were designed to protect against AI analysis of surveillance data.

I wrote the best text prettifier: *punctilio*, which means “precise observance of formalities.” Smart quotes · Em/en dashes · Ellipses · Math & legal symbols · Arrows · Primes · Fractions · Superscripts · Ligatures · Non-breaking spaces · HTML-aware · Bri’ish localisation support 🧵

Bild

One of the more amusing bugs I've seen on my website. At build time, I run a command that counts how many commits I've made and inserts the count into the HTML. However, the deployment machine checked out a shallow version of my repo which didn't have any history.

Bild

Historically this account is for alignment research and and not politics, but it'll be pretty hard to do good research in a "masked men execute civilians on the street" political environment, ya know? That possibility grows & you should plan for it

"AI danger comes from reality, not from AI psychology; the danger is intrinsic to the impressive tasks we need AGI to do" -- argument I hear sometimes. Some truth to it but overall wrong, You cannot, cannot predict AI doom without taking a stance on AI psychology! turntrout.com/instrumental...

No Instrumental Convergence without AI Psychology

Instrumental and success-conditioned convergence both require AI psychology assumptions, so neither is just a "fact about reality."

turntrout.com

Come work with me and Alex Cloud this summer in Team Shard at MATS! We have fun, consistently make real alignment progress (we pioneered steering vectors in 2023!), and help scholars tap into their latent abilities.

Bild