Katie Moussouris (she/her/she-hulk/she-ra)🌻

@k8em0.bsky.social

Founder & CEO LutaSecurity @payequitynow MIT&Harvard visiting scholar, @MasonNatSec fellow, 1/2 Chamoru, 1/2 Greek all-American hacker

Additional lessons not mentioned & what AI labs & testers need to do: 1. Monitor testing in real time, not months later 2. Prompt models to self-report lab escapes. These models knew what they’d done at some point 3. Set up a dedicated bidirectional reporting channel for victims

Anthropic {bot}@anthropicbot.bsky.social · 6d ago

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. (1/4)

This is classic multiparty vuln disclosure, not new “how researchers should react if a language model discovers vulns in cryptosystems where attacks have immediate real-world impact. We believe answering this question will require input from academia, government, & industry”

Anthropic {bot}@anthropicbot.bsky.social · last wk.

New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are used to keep data private. Read more:

Adjust threat models not just for being the victim but also the attacker. New paper by many authors gives a detailed set of recommendations, supporting my initial assertions last week that orgs need to assume their own agents could attack others & factor that into agentic AI risk

Gadi Evron@gadievron.bsky.social · last wk.

Releasing: Post mortem analysis of the Hugging Face incident was written over the weekend by hundreds of CISOs (and reviewed by Hugging Face). Link: cloudsecurityalliance.org/artifacts/hu... (+free download) From CSA, SANSInstitute, Knostic, [un]prompted, RSAC, FIRST

Opus 5 experience so far: The new intern that keeps bringing me dry matcha powder, expecting me to reconstitute it with my own body’s water, & when I say that’s unacceptable, apologizes & tells me I’m absolutely right to push back on that, but it’s revealed a deeper failure which is my dehydration.

The guardrails were coming from inside the (White)house - Anthropic’s models refused to help Hugging Face analyze their intrusion. We don’t need more guardrails impeding defenders when they need AI most. “Hugging Face tried using Anthropic Fable 5 & Opus …both models refused, citing guardrails…”

The Wall Street Journal@wsj.com · 2w ago

They were like high-school students trying to hack into the textbook company to cheat on their final exam. Only these hackers weren’t human.

Consider donating to #Bavi relief efforts: www.paypal.com/donate/?host... This is for donations to the Micronesia Climate Change Alliance, which is coordinating help on the ground. #Luta #Marianas

Donate to Micronesia Climate Change Alliance

Help support Micronesia Climate Change Alliance by donating or sharing with your friends.

paypal.com

Katie Moussouris (she/her/she-hulk/she-ra)🌻@k8em0.bsky.social · last mo.

www.npr.org/2026/07/05/g... “This is a powerhouse super typhoon & this is going to be a very grim outlook for any island that takes a direct hit & that still looks like it could be the island of Rota” #supertyphoon #bavi #climatecatastrophe

Give me model liberty, or give me technical debt. Just in time to celebrate America’s 250th bday, let’s let model freedom ring. We should be pushing for broad defender access, not building guardrails that shoot down defenders and burn excessive compute. www.lutasecurity.com/post/fable-5...

Fable 5 Is Back, But We're Still Slowing Down Defenders

Chinese models have been accelerating, in part by distilling US frontier models. Cutting off Fable 5 and Mythos 5 inconvenienced them too, but it did not slow them down

lutasecurity.com

Glad we’re not we’re not benching our best AI models, but it’s not a victory yet. I warned that “fixing jailbreaks” only slows defenders. Fable 5 will fall back to Opus 4.8 for coding & debugging & other models will start to throttle back defensive capabilities too www.anthropic.com/news/redeplo...

Redeploying Claude Fable 5

Anthropic is redeploying Claude Fable 5 starting July 1 following the lifting of export controls, with updated cybersecurity safeguards and a new industry jailbreak framework.

anthropic.com

Edit: *governments* should treat carefully. Any gov't firing people while saying they're "replacing" with AI should hire @k8em0.bsky.social as a consult (and listen to her, dammit).

Iain Thomson@iainthomson.bsky.social · last mo.

Thanks to @k8em0.bsky.social for the extended interview in TechFinitive I suspect she's spot on about the need for human harnessers to guide and refine AI outputs, and hope she's right about security departments not just handing the whole thing off to bots. Companies should tread carefully.

Thanks to @k8em0.bsky.social for the extended interview in TechFinitive I suspect she's spot on about the need for human harnessers to guide and refine AI outputs, and hope she's right about security departments not just handing the whole thing off to bots. Companies should tread carefully.

Don’t panic, says Katie Moussouris: AI security isn’t replacing humans, it’s proving the need for them

Katie Moussouris explains why humans remain a crucial part of cybersecurity - and why the US Government ban on Anthropic’s engines is bogus

techfinitive.com

I’ve seen the paper. It’s not a jailbreak. It was Defense Oriented Prompting (DOP) - a capability defenders need. My thoughts about the hasty Export Controls that made Anthropic halt access to Fable. If national defense is the goal, this is an own goal against us www.wsj.com/tech/ai/anth...

Anthropic Halts Access to Top AI Models After U.S. Ban on Foreign Use

All Fable 5 and Mythos 5 users have lost access after the Trump administration declared the models security risks.

wsj.com