Hadas Orgad

@hadasorgad.bsky.social

📢 We’re looking for reviewers for the Actionable Interpretability workshop @ActInterp ! If you’re interested in helping review submitted papers, please sign up here: forms.gle/7pihaQuSQ2Wq... Your expertise would be greatly appreciated!

Reviewer Form - Actionable Interpretability Workshop

This form collects information on reviewers for the workshop Actionable Interpretability @ COLM 2026. Please take note of the details for the review process: Important Dates: Review Start: June 25...

forms.gle

New paper: LLMs encode harmful content generation in a distinct, unified mechanism Using weight pruning, we find that harmful generation depends on a tiny subset of the weights that are shared across harm types and separate from benign capabilities. 🧵

Bild

Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵

Bild

Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!

Logo for MIB: A Mechanistic Interpretability Benchmark