Alex Irpan

@alexirpan.bsky.social

Research Scientist @ Google DeepMind. Formerly Robotics, now AI Safety. Has a blog. Views are my own.

In retrospect "AI CEOs become celebrities and attract corresponding levels of desired and undesired attention" was a very obvious prediction that many people failed to account for (including me)

1. Obviously terrible to have a Molotov thrown against your house, not appropriate response. 2. Of all analogies to make, "ring of power" is a choice, given the story's theme that the only way to stop the ring's destructive power is to destroy it. blog.samaltman.com/2279512

-

Here is a photo of my family. I love them more than anything. Images have power, I hope. Normally we try to be pretty private, but in this case I am sharing a photo in the...

blog.samaltman.com

First paper since switching into AI safety team🎉 We look at problems that could be solved if the model behaved consistently over a set of prompts, and tried training that in output space and internal activations. Both were effective. See thread or paper for details.

Alex Turner@turntrout.bsky.social · 9mo ago

New Google DeepMind paper: "Consistency Training Helps Stop Sycophancy and Jailbreaks" by @alexirpan.bsky.social, me, Mark Kurzeja, David Elson, and Rohin Shah. (thread)

The abstract of the consistency training paper.

"I don't play gacha games because they're a scam" vs "Let me do one more hyperparam sweep before giving up. One more prompt tuning run. I swear we'll beat baseline. I know it's gonna beat the baseline this time. It's gonna win. This time for sure."

The ship has sailed, but I wish the ML reporting default was % incorrect rather than % correct. It better matches loss curves and magnifies the capture of edge cases. 95% accuracy -> 97.5% accuracy = meh 5% error -> 2.5% error = omg we've halved the error rate