Cas (Stephen Casper)

@scasper.bsky.social

Computer scientist working on AI safeguards and governance research. Assistant professor @harvardkennedy.bsky.social @harvard.edu. https://stephencasper.com/

None of the 19 incidents that UK AISI found were from 'helpful-only' models. There is a case to be made that setting Mythos- or GPT 5.6-Sol-level AI cyberagents to run without a robust real-time monitoring setup is an inherently (perhaps abnormally) dangerous activity.

BildBild

“The agents went rogue.” Kinda, but imagine that a zoo had a habit of not shutting the door on animal enclosures or putting elephants behind chicken wire — all with no zookeepers in sight. The fault is not in our stars.

🧵 If trends hold, expect a Mythos-level open-weight model around Christmas. Meanwhile, open models are important but also a massive hole in most agendas for safe AI. Here's a 🧵 of my thoughts & research agenda on the technical & political challenges we need to address.

Bild

I was wondering if there was any research about how AI companies sometimes perversely keep their safety research to themselves to build a moat around it and gain a competitive edge over competitors. I found one. It's a cool paper. Sharing here in case anyone's interested.

Bild

🧵 In AI, we are used to seeing graphs that start to exhibit hockey stick behavior around 2023-2025. But that's a little bit funny and incongruous in light of how relatively little the Overton window has changed with AI lawmaking since 2024...

Bild

The summary released today of the FRONTIER Act is cool. It seems like a pretty rigorous bill. Based on the summary, in my opinion, it might be good enough to be worth passing. But I would still tweak a few things. Here is a brainstorm of 8 ideas. 🧵

Bild

Want to get up to speed on what researchers have been saying about internal deployment of AI lately? I would recommend checking out these four papers. Let me know in the replies if I'm missing something.

BildBildBildBild

OpenAI's internally deployed models hacking Hugging Face does not seem to have been unpredictable or inevitable. We talked about the root of the problem & what policymakers can do about it back in February. Props to Joe for hitting the nail on the head.

BildBildBild

🚨 New paper: Some, but not all, AI companies make corporately-loyal models. xAI, DeepSeek, Anthropic, & OpenAI models all downplay company controversies. Google, Meta, & Alibaba models don't. The findings are clear, but we are pretty confused as to why... 🧵 @finke.dev

Bild

Stability is now being sued (alongside xAI) for abetting the production of AI NCII/CSAM due to how it developed & released several open-weight models. Anyone interested in whether AI companies will be held liable for foreseeable, mitigatable *downstream* harms should follow this.

BildBildBildBild

Just saw this new paper. It was already known that models from Stability and Alibaba dominate the image & video NCII ecosystems, respectively, but I didn't know they were *this* dominant. Just a few socially reckless companies are the principal enablers of AI NCII abuse.

Bild

Lennart Finke and I will release a paper on Monday about how some AI developers tend to make models that differentially downplay company controversies. Below (🧵) is a link to a 1-question Google form for you to guess the results before they're out. (They might surprise you.)

Bild

I just gave Bernie's AI Wealth Fund Act a close read. The wealth redistribution via this Act would be enormous. But it does something else far more impactful... 🧵Here's what it does, plus 3 things I'd change -- one of which I think, if unaddressed, could be a fatal flaw.

BildBild

Yes, countries CAN cooperate on AI cyber risks. Countries like China and the US love to constantly cyberattack each other. Because of this, I have heard a few people (under Chatham House rules) speculate that cyberdefense is an AI risk domain in which international cooperation is unlikely...

If I were Anthropic, I would honestly be overjoyed at the Trump admin blocking Fable. - It's probably temporary - It's free publicity - It distinguishes Anthropic w.r.t. other companies - People want what they can't have - The admin doesn't have much credibility anyway

A generationally important US House primary for AI and tech policy is happening on June 23 in New York's 12th district. If you live in NY-12, and if you believe that AI safeguards, transparency, and accountability are critical, I hope you consider voting for @alexbores.nyc (D).

BildBild

Now that I have your attention by posting this spinning point cloud GIF, I'd like to propose a litmus test for AI mechanistic interpretability research. You might call it the "interp hammer" test...🧵

Anthropic and OpenAI are publicly pointing out how having the option to slow down AI would offer a potentially critical form of optionality in the future. The correct response for any policymaker should be "Damn, this is serious. How can I help build that capacity?"

BildBild