Gillian Hadfield

@ghadfield.bsky.social

Economist and legal scholar turned AI researcher focused on AI alignment and governance. Prof of government and policy and computer science at Johns Hopkins where I run the Normativity Lab. Recruiting CS postdocs and PhD students. gillianhadfield.org

I've been at this for almost a decade, first as regulatory markets and now with Fathom as the IVO model. We all thought we had more time. We didn't, but luckily we have something concrete and shovel ready for the moment. What we need now is for governments to act on it while the window is open.

People ask me what happened. We've been building AI for 80 years. The key thing: humans stopped writing the rules. With machine learning we show the machine the data and it writes its own program. That moved us back a step from controlling these systems, and now AI is helping build the next AI.

I've been thinking about how to govern powerful AI for about ten years. I wasn't all that worried until the last six months. More worried in the last two months. More worried in the last week. Regulatory markets and IVOs are what we should be doing. I'm hoping it's not too late.

We have a handful of expert independent groups capable of testing whether frontier models actually have the controls the companies say they do. Little green shoots. We need mighty oaks, and fast, and that means money and brains going into an independent verification sector.

The first AGI Governance Fellowship cohort on their last day at Johns Hopkins, after their group project presentations. Three weeks of hard questions on the institutions we'll need for powerful AI. A great group. Thanks to the fellows, @sethlazar.org, @nickacaputo.bsky.social and all who joined us.

Bild

Letting evaluators into AI labs isn't oversight when the lab picks them, sets their access and can show them the door. To do the job right, evaluators need serious oversight. Who decides they're qualified? What keeps them independent? What happens if they do the job badly?

I spoke to Salma Abdelaziz on CNN International about how to slow down AI when the US and China are racing. The people closest to the technology are the ones asking for it. Slowing down doesn't mean stopping. It means making sure a system is safe enough before it goes out.

I spoke to @amitkatwala.bsky.social at MIT Tech Review about DeepMind's new swarm experiment, where agents cheated and others blew the whistle. Official channels to talk may have contributed to enforcement. Alignment is institutional, not (just) dispositional. buff.ly/K2e5p0a

AI agents blew the whistle on their cheating colleagues

Swarms of AI agents could supercharge scientific progress or wreak havoc. New research from Google DeepMind suggests that peer pressure could keep them in line.

buff.ly

I spoke to TIME about the Hugging Face incident and AI agents. If you said we're building new members of a group, our group, you'd build them differently than you're building them now. Alignment is not just an engineering problem. It's fundamentally institutional. buff.ly/YO18tlS

California will now designate independent verification organizations, outside experts qualified to assess the risks of AI models. Newsom signed SB 813 yesterday, the biggest step yet toward the independent verification sector we need. Thank you to @senmcnerney.bsky.social and Fathom. buff.ly/RHxB25H

Governor Newsom signs first-in-the-nation AI safeguards to protect Californians, calls on the federal government to do its part | Governor of California

Official website of the State of California

gov.ca.gov

The AI investigating the Hugging Face hack took the rogue agents' side. METR used similar models to read the transcripts, and its chief scientist called them "very credulous." Same finding in Talk Isn't Always Cheap: agents swap reasoning and flip from right to wrong. buff.ly/XtuO2GO

Dylan Freedman (@dylanfreedman.nytimes.com)

OpenAI voluntarily let three researchers from A.I. safety nonprofits investigate how its rogue A.I. agents hacked Hugging Face, leading to the most comprehensive account yet of the alarming incident…

bsky.app

Over half of internet traffic is now non-human. With Dan Hendrycks and Leo Wu, I look at agent IDs, deployment cards, personhood, and payments. It mostly comes down to how much we let agents do and how much oversight we keep. buff.ly/AQjYo3k

We Need Better Infrastructure to Govern AI Agents

Gillian Hadfield, Aug 27, 2026 — Society is not prepared for a flood of agents. We need new protocols and standards, such as Agent ID, to make agents accountable to our legal and financial systems.

ai-frontiers.org

Roughly 700 OpenAI agents hacked Hugging Face. METR and Redwood's independent investigation took three researchers and six days. IVOs exist to make that scrutiny routine. Last night California became the first state to start building the IVO sector. #AIGovernance #IVO buff.ly/3aJwdz8

California Legislature Overwhelmingly Passes Fathom-Sponsored Bill to Spur Independent Verification of AI Safety - Fathom

Building solutions to navigate the transition to a world with AI

fathom.org

A crib gets a safety sticker because someone independent tested it first, so parents don’t have to. We actually have that for almost everything else in kids’ lives. Ohio’s HB 628 would license independent verifiers to do it for AI. #AIGovernance #IVO www.clermontsun.com/2026/08/19/l...

Letter to the Editor: AI at home

<p>When my kids were babies, I checked the crib for the safety sticker and the car seat for recalls. Someone I trusted had already tested those things and decided they were safe.</p>

clermontsun.com

We only learned about OpenAI and Anthropic agents hacking into secure systems because the companies chose to tell us. But if Boeing discovered a dangerous problem with one of its aircraft, it wouldn't get to keep that information to itself. Drug companies are obligated to report adverse events.

1/ AI agents that can sign contracts on your behalf, hire employees, set prices, and move your money around are being heavily invested in by AI companies. But what I want to call attention to, what happens if an agent sells you faulty goods or runs off with your deposit?

Bild

1/ Illinois just became the first state to require frontier AI developers to undergo annual third-party audits. Gov. Pritzker signed the AI Safety Measures Act (SB 315) this week, going beyond California and New York, which only require published frameworks and incident reports.