Stella Biderman

@stellaathena.bsky.social

I make sure that OpenAI et al. aren't the only people who are able to study large scale AI systems.

These concerns are very much not hypothetical. In 2024 a private benchmark called FrontierMath developed by a third party called Epoch came out. We later learned that OpenAI funded the creation of the benchmark and while nobody else had the test Qs, OAI did. ai-frontiers.org/articles/don...

Don’t Let AI Developers Hire Their Own Referees | AI Frontiers

Gabriel Weil, Jul 29, 2026 — Letting AI developers pick their own safety auditors creates a conflict of interest. Requiring liability insurance instead would put insurers’ own capital behind risk asse...

ai-frontiers.org

If you believe their story, OpenAI accidentally committed cyber warfare against Hugging Face while doing what they thought was an internal test of a model without internet access. In six months time will be obvious they’re faced effectively no sanction for this.

When I was in middle school (late 2000s), I was very interested in mathematics. I thought that the validity of the proof of the 4-color theorem was disputed in mathematics and was shocked when I learned about FLT, but I knew that humans would never beat an AI at chess ever again

The ability of tech co. to produce "intelligent" systems that are incompetent at doing any of the tasks I actually want them to do is mind-boggling. TIL that ChatGPT and Claude generally don't agree when you give them two papers and ask how many citations they have in common.

In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵

Bild

I had given Anthropic a lot of credit for turning down the DoD and its trillions of dollars, especially as it seemed to be the only example of any AI company making any financial sacrifices for moral principles. Of course, it turns out to be not really true. www.axios.com/2026/04/19/n...

Scoop: NSA using Anthropic's Mythos despite Defense Department blacklist

The government's cybersecurity needs are outweighing the Pentagon's feud with Anthropic.

axios.com

Excited to be on my way to @iclr-conf.bsky.social! Come stop by our posters and hit me up. I'm especially excited to talk about - Open weight safety - Training dynamics and interpretability over time - Memorization and machine unlearning - Open data - Rigorous experimental design

@eleutherai.bsky.social · 4mo ago

Looking for EleutherAI @iclr-conf.bsky.social? Come by our posters! If you're in our discord, we have a thread #general > ICLR 2026 Meetup you can join to coordinate with @stellaathena.bsky.social, Goncalo Paulo, @norabelrose.bsky.social, and members of our community who will be there!

EleutherAI at ICLR 2026 — where to find our work. Apr 23–27, 2026 in Rio de Janeiro, Brazil. 3 main-track papers (50% acceptance, 3 of 6 submitted) and 1 workshop paper (100%, 1 of 1).
Main track posters, all at Pavilion 4, 10:30 AM – 1:00 PM: (1) "Sparse Autoencoders Trained on the Same Data Learn Different Features" by Gonçalo Paulo and Nora Belrose — Thu Apr 23, board P4-#4004. (2) "Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs" by Kyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman — Fri Apr 24, board P4-#4115. (3) "Evaluating SAE Interpretability without Generating Explanations" by Gonçalo Paulo and Nora Belrose — Sat Apr 25, board P4-#4007.
Workshop paper at the ICBINB Workshop: "Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA" by Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj, Patrick Liu, Timothy Chung, Drishti Sharma, Akshata A., Kranthi Kiran, Wesley Tam, and Bala Krishna S Vegesna — Mon Apr 27, 13:00–14:25, Room 201C.

FISA 207 is blatantly illegal and immoral and has always been obviously so. Republicans are pretending to not know this, just like Democrats did during the Biden administration. This is bipartisan evil.