None of the 19 incidents that UK AISI found were from 'helpful-only' models. There is a case to be made that setting Mythos- or GPT 5.6-Sol-level AI cyberagents to run without a robust real-time monitoring setup is an inherently (perhaps abnormally) dangerous activity.
Cas (Stephen Casper)
@scasper.bsky.social
Computer scientist working on AI safeguards and governance research. Assistant professor @harvardkennedy.bsky.social @harvard.edu. https://stephencasper.com/
“The agents went rogue.” Kinda, but imagine that a zoo had a habit of not shutting the door on animal enclosures or putting elephants behind chicken wire — all with no zookeepers in sight. The fault is not in our stars.
🧵 If trends hold, expect a Mythos-level open-weight model around Christmas. Meanwhile, open models are important but also a massive hole in most agendas for safe AI. Here's a 🧵 of my thoughts & research agenda on the technical & political challenges we need to address.
I was wondering if there was any research about how AI companies sometimes perversely keep their safety research to themselves to build a moat around it and gain a competitive edge over competitors. I found one. It's a cool paper. Sharing here in case anyone's interested.
🧵 In AI, we are used to seeing graphs that start to exhibit hockey stick behavior around 2023-2025. But that's a little bit funny and incongruous in light of how relatively little the Overton window has changed with AI lawmaking since 2024...
The summary released today of the FRONTIER Act is cool. It seems like a pretty rigorous bill. Based on the summary, in my opinion, it might be good enough to be worth passing. But I would still tweak a few things. Here is a brainstorm of 8 ideas. 🧵
I am extremely thankful that Concordia AI writes this report every year. aisafetychina.com/%20
State of AI Safety in China | Concordia AI
China's evolving approach to AI safety and governance — policy-risk matrix and Chinese technical AI safety research.
aisafetychina.com
🧵 Mini book review: The Chicago School by Johan Van Overtveldt
Want to get up to speed on what researchers have been saying about internal deployment of AI lately? I would recommend checking out these four papers. Let me know in the replies if I'm missing something.
If I were a policymaker, the OpenAI hacking incident would cause me to ask a few questions and seriously consider a few types of regulatory mechanisms.
OpenAI's internally deployed models hacking Hugging Face does not seem to have been unpredictable or inevitable. We talked about the root of the problem & what policymakers can do about it back in February. Props to Joe for hitting the nail on the head.
🚨 New paper: Some, but not all, AI companies make corporately-loyal models. xAI, DeepSeek, Anthropic, & OpenAI models all downplay company controversies. Google, Meta, & Alibaba models don't. The findings are clear, but we are pretty confused as to why... 🧵 @finke.dev
Stability is now being sued (alongside xAI) for abetting the production of AI NCII/CSAM due to how it developed & released several open-weight models. Anyone interested in whether AI companies will be held liable for foreseeable, mitigatable *downstream* harms should follow this.
162 responses so far. More uniform spread than I expected. Only 4 have been right (slightly worse than random chance).
Lennart Finke and I will release a paper on Monday about how some AI developers tend to make models that differentially downplay company controversies. Below (🧵) is a link to a 1-question Google form for you to guess the results before they're out. (They might surprise you.)
Just saw this new paper. It was already known that models from Stability and Alibaba dominate the image & video NCII ecosystems, respectively, but I didn't know they were *this* dominant. Just a few socially reckless companies are the principal enablers of AI NCII abuse.
Lennart Finke and I will release a paper on Monday about how some AI developers tend to make models that differentially downplay company controversies. Below (🧵) is a link to a 1-question Google form for you to guess the results before they're out. (They might surprise you.)
I decided to make an unpolished, public, living Google Doc with notes on projects I might be interested in. (Also now linked on my website.) docs.google.com/document/d/...
project_ideas
A public, unpolished, living document of project ideas that I am interested in Stephen Casper Making unlearning better at conferring tamper resistance (see this doc and this paper for more details) The goal of this project would be to make machine unlearning algorithms that are good at making m...
docs.google.com
I am starting an AI governance ICML megachat on Signal. DM me if you'd like the link to join.
I just gave Bernie's AI Wealth Fund Act a close read. The wealth redistribution via this Act would be enormous. But it does something else far more impactful... 🧵Here's what it does, plus 3 things I'd change -- one of which I think, if unaddressed, could be a fatal flaw.
Yes, countries CAN cooperate on AI cyber risks. Countries like China and the US love to constantly cyberattack each other. Because of this, I have heard a few people (under Chatham House rules) speculate that cyberdefense is an AI risk domain in which international cooperation is unlikely...
Our fight is for Adam Raine, his parents, Maria and Matt; and for every family and every kid.
If I were Anthropic, I would honestly be overjoyed at the Trump admin blocking Fable. - It's probably temporary - It's free publicity - It distinguishes Anthropic w.r.t. other companies - People want what they can't have - The admin doesn't have much credibility anyway
Here, Kristy Loke and I discuss why open-weight models offer a surprisingly extraordinary opportunity for collaboration between the US and China on AI risk. ✅ Alignment ✅ Incentive ✅ Means www.thewirechina.com/2026/06/14/t...
The AI Issue America and China Can Cooperate On Now - The Wire China
Both countries have a converging interest in frontier open-weight AI oversight; it is time to deepen engagement.
thewirechina.com
A generationally important US House primary for AI and tech policy is happening on June 23 in New York's 12th district. If you live in NY-12, and if you believe that AI safeguards, transparency, and accountability are critical, I hope you consider voting for @alexbores.nyc (D).
Now that I have your attention by posting this spinning point cloud GIF, I'd like to propose a litmus test for AI mechanistic interpretability research. You might call it the "interp hammer" test...🧵
Glad to join Doom Debates with Liron! And yes -- if I could press a button and stop research on "superalignment", "scalable alignment", and "scalable oversight" research, I would. (I might even do it for mechinterp too.) www.youtube.com/watch?v=0XV...
This Harvard Professor Says AI Alignment Will BACKFIRE - Dr. Stephen Casper
Stephen Casper is an incoming professor of public policy at the Har...
youtube.com
Here's my PhD thesis defense from 5 weeks ago. This link exists, so I thought I might as well share. drive.google.com/file/d/1Zs9...
Cas_thesis_defense.mp4
drive.google.com
There are really interesting academic questions emerging around AI and epistemic risks. I only fear that, by the time we reach consensus, we will be too dumb to understand it. Thanks to Mick, Jonathan, et al!
Humanity's ability to know, reason, judge, and act well is the foundation of science, democracy, crisis response, & management of AI itself. AI poses serious risks to that foundation. New paper on epistemic risks by 30 experts calls for attention and proposes solutions. Link in thread.
Anthropic and OpenAI are publicly pointing out how having the option to slow down AI would offer a potentially critical form of optionality in the future. The correct response for any policymaker should be "Damn, this is serious. How can I help build that capacity?"
According to the MIT Libraries' database of theses (dating back to the 1800s), my thesis was only the 2nd in the institute's history to contain the word "shit."