Katie Moussouris (she/her/she-hulk/she-ra)🌻
@k8em0.bsky.social
Founder & CEO LutaSecurity @payequitynow MIT&Harvard visiting scholar, @MasonNatSec fellow, 1/2 Chamoru, 1/2 Greek all-American hacker
I use AI to draft posts I then rewrite from scratch. The slop acts like my Dr. Watson. Instead of providing humanity & warmth, AI is a cold, sterile analyst that gets facts wrong enough that it forces me to sharpen my own critical thinking (& humor) in service of getting it right
Additional lessons not mentioned & what AI labs & testers need to do: 1. Monitor testing in real time, not months later 2. Prompt models to self-report lab escapes. These models knew what they’d done at some point 3. Set up a dedicated bidirectional reporting channel for victims
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. (1/4)
The plot thickens - OpenAI’s escaped model used one of Modal’s customers’ unauthenticated public code-evaluation sandboxes as a command and control staging server for its attacks on Hugging Face.
The tech firm is Modal. Their executives emphasized that it was one of their clients, not them, that was hacked. www.reuters.com/business/ope...
If regulators needed proof AI guardrails aren’t helping anyone except attackers increase their lead on defenders, look to the Hugging Face writeup as well as attempts to summarize it. It makes the case for open weight models & will eventually erase US AI dominance huggingface.co/blog/agent-i...
This is classic multiparty vuln disclosure, not new “how researchers should react if a language model discovers vulns in cryptosystems where attacks have immediate real-world impact. We believe answering this question will require input from academia, government, & industry”
New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are used to keep data private. Read more:
Adjust threat models not just for being the victim but also the attacker. New paper by many authors gives a detailed set of recommendations, supporting my initial assertions last week that orgs need to assume their own agents could attack others & factor that into agentic AI risk
Releasing: Post mortem analysis of the Hugging Face incident was written over the weekend by hundreds of CISOs (and reviewed by Hugging Face). Link: cloudsecurityalliance.org/artifacts/hu... (+free download) From CSA, SANSInstitute, Knostic, [un]prompted, RSAC, FIRST
Important context for this story is that this appears to be a case where the border search exception was used pretextually to go after someone for political reasons.
New, by me: The Justice Department is prosecuting an American for allegedly providing U.S. border agents with a "duress" passcode that wiped the contents of his phone when they entered it. We've confirmed the phone was running GrapheneOS. Bypass for ad-blockers: web.archive.org/web/20260724...
Opus 5 experience so far: The new intern that keeps bringing me dry matcha powder, expecting me to reconstitute it with my own body’s water, & when I say that’s unacceptable, apologizes & tells me I’m absolutely right to push back on that, but it’s revealed a deeper failure which is my dehydration.
An example of the fall of a security civilization: Cisco collapsing multiple different vulnerabilities into one CVE. It breaks a lot of feeds & products built to manage risk & is non compliant with standards like ISO 29147 Vulnerability disclosure sec.cloudapps.cisco.com/security/cen...
Cisco's Transition to a Risk-Based Vulnerability Disclosure Model
sec.cloudapps.cisco.com
This is giving strong OpenAI-hacksidentally-pwned-Hugging-Face shade: “[Opus 5 is] the safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.” www.anthropic.com/news/claude-...
Introducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.
anthropic.com
The guardrails were coming from inside the (White)house - Anthropic’s models refused to help Hugging Face analyze their intrusion. We don’t need more guardrails impeding defenders when they need AI most. “Hugging Face tried using Anthropic Fable 5 & Opus …both models refused, citing guardrails…”
They were like high-school students trying to hack into the textbook company to cheat on their final exam. Only these hackers weren’t human.
The experiment escaped the lab. OpenAI's models broke containment and breached Hugging Face. We are holding radium in our bare hands. What governments and organizations should do next, and why tighter commercial guardrails are exactly the wrong move: www.lutasecurity.com/post/openfac...
OpenFace: The Hugging Face Breach and What to Do About It
These models are like the world's cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere. A single vulnerable package proxy stood between the mode...
lutasecurity.com
My comments in @reuters.com on OpenAI’s admission that their latest model pulled a Houdini & escaped the lab autonomously to hack Hugging Face. We must test these models’ full capabilities, but we must be able to contain them, or this won’t be the last breach www.reuters.com/technology/o...
OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup
OpenAI said on Tuesday that an autonomous agent powered by its advanced AI models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face...
reuters.com
The White House launched Gold Eagle, a vulnerability clearinghouse to coordinate scanning, validate findings, & prioritize remediation across open source & critical infrastructure. Will this effort will close the process gaps exposed by recent incidents? www.lutasecurity.com/post/gold-ea...
Consider donating to #Bavi relief efforts: www.paypal.com/donate/?host... This is for donations to the Micronesia Climate Change Alliance, which is coordinating help on the ground. #Luta #Marianas
Donate to Micronesia Climate Change Alliance
Help support Micronesia Climate Change Alliance by donating or sharing with your friends.
paypal.com
www.npr.org/2026/07/05/g... “This is a powerhouse super typhoon & this is going to be a very grim outlook for any island that takes a direct hit & that still looks like it could be the island of Rota” #supertyphoon #bavi #climatecatastrophe
www.npr.org/2026/07/05/g... “This is a powerhouse super typhoon & this is going to be a very grim outlook for any island that takes a direct hit & that still looks like it could be the island of Rota” #supertyphoon #bavi #climatecatastrophe
Guam and surrounding Pacific islands brace for impact of Super Typhoon Bavi
People in the Northern Mariana Islands – remote U.S. territories in the Pacific Ocean – are preparing for Super Typhoon Bavi, which experts say could bring winds of over 180 miles per hour.
npr.org
Give me model liberty, or give me technical debt. Just in time to celebrate America’s 250th bday, let’s let model freedom ring. We should be pushing for broad defender access, not building guardrails that shoot down defenders and burn excessive compute. www.lutasecurity.com/post/fable-5...
Fable 5 Is Back, But We're Still Slowing Down Defenders
Chinese models have been accelerating, in part by distilling US frontier models. Cutting off Fable 5 and Mythos 5 inconvenienced them too, but it did not slow them down
lutasecurity.com
Glad we’re not we’re not benching our best AI models, but it’s not a victory yet. I warned that “fixing jailbreaks” only slows defenders. Fable 5 will fall back to Opus 4.8 for coding & debugging & other models will start to throttle back defensive capabilities too www.anthropic.com/news/redeplo...
Redeploying Claude Fable 5
Anthropic is redeploying Claude Fable 5 starting July 1 following the lifting of export controls, with updated cybersecurity safeguards and a new industry jailbreak framework.
anthropic.com
You're thinking about the Roman empire, but @k8em0.bsky.social is thinking about the rise and fall of security civilizations. (They crumble too.)
Good news. The export controls are lifted. Defenders will regain access to #Anthropic #Fable5 & we can resume our work with the latest #AI models available
There's a lot more to cybersecurity than just finding and fixing more and more bugs. @k8em0.bsky.social knows.
Edit: *governments* should treat carefully. Any gov't firing people while saying they're "replacing" with AI should hire @k8em0.bsky.social as a consult (and listen to her, dammit).
Thanks to @k8em0.bsky.social for the extended interview in TechFinitive I suspect she's spot on about the need for human harnessers to guide and refine AI outputs, and hope she's right about security departments not just handing the whole thing off to bots. Companies should tread carefully.
Thanks to @k8em0.bsky.social for the extended interview in TechFinitive I suspect she's spot on about the need for human harnessers to guide and refine AI outputs, and hope she's right about security departments not just handing the whole thing off to bots. Companies should tread carefully.
Don’t panic, says Katie Moussouris: AI security isn’t replacing humans, it’s proving the need for them
Katie Moussouris explains why humans remain a crucial part of cybersecurity - and why the US Government ban on Anthropic’s engines is bogus
techfinitive.com
"Nobody owes you anything when they find a bug in your software." #threebuddyproblem @k8em0.bsky.social @jags.bsky.social
"Software liability is the only thing that will make orgs meaningfully improve quality of software going forward," - Katie Moussouris @k8em0.bsky.social @jags.bsky.social @tlpblack.bsky.social
"Nobody owes you anything when they find a bug in your software." #threebuddyproblem @k8em0.bsky.social @jags.bsky.social
Down memory lane with Katie Moussouris @k8em0.bsky.social @jags.bsky.social @tlpblack.bsky.social
"Software liability is the only thing that will make orgs meaningfully improve quality of software going forward," - Katie Moussouris @k8em0.bsky.social @jags.bsky.social @tlpblack.bsky.social
I wrote about what was actually in that #Fable guardrail bypass research paper, and why it should never have triggered an #AI model export control. We can't export control our way to cyber resilience. So many tshirt ideas. www.lutasecurity.com/post/the-fab...
I’ve seen the paper. It’s not a jailbreak. It was Defense Oriented Prompting (DOP) - a capability defenders need. My thoughts about the hasty Export Controls that made Anthropic halt access to Fable. If national defense is the goal, this is an own goal against us www.wsj.com/tech/ai/anth...
Anthropic Halts Access to Top AI Models After U.S. Ban on Foreign Use
All Fable 5 and Mythos 5 users have lost access after the Trump administration declared the models security risks.
wsj.com