📢 We’re looking for reviewers for the Actionable Interpretability workshop @actinterp.bsky.social! If you’re interested in helping review submitted papers, please sign up here: forms.gle/VpLJpkM6zw3V... Your expertise would be greatly appreciated!
Sarah Wiegreffe
@sarah-nlp.bsky.social
Research in NLP (mostly LM interpretability & explainability). Assistant prof at UMD CS + CLIP. Previously @ai2.bsky.social @uwnlp.bsky.social Views my own. sarahwie.github.io
Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵
Come join TRAILS as a postdoc at UMD (and work w folks at GW, MSU & Cornell) to conduct research and scholarship focused on approaches to AI that advance trust and trustworthiness with a great group of colleagues! 🌐 go.umd.edu/trails-postd... 🗓️ Summer/Fall 2026 start
TRAILS UMD Post Doctoral Associate Job Description - Spring 2026
Post Doctoral Associate Institute for Trustworthy AI in Law & Society February 2026 The Institute for Trustworthy AI in Law & Society (TRAILS) and the University of Maryland aim to transform the pr...
go.umd.edu
If you're at #ICML2025, chat with me, @sarah-nlp.bsky.social, Atticus, and others at our poster 11am - 1:30pm at East #1205! We're establishing a 𝗠echanistic 𝗜nterpretability 𝗕enchmark. We're planning to keep this a living benchmark; come by and share your ideas/hot takes!
I am at #ICML2025! 🇨🇦🏞️ Catch me: 1️⃣ Presenting this paper👇 tomorrow 11am-1:30pm at East #1205 2️⃣ At the Actionable Interpretability @actinterp.bsky.social workshop on Saturday in East Ballroom A (I’m an organizer!)
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
This week is #ICML in Vancouver, and a number of our researchers are participating. Here's the full list of Ai2's conference engagements—we look forward to connecting with fellow attendees. 👋
A bit late to announce, but I’m excited to share that I'll be starting as an assistant professor at UMD CS @univofmaryland.bsky.social this August. I'll be recruiting PhD students this upcoming cycle for fall 2026. (And if you're a UMD grad student, sign up for my fall seminar!)
🚨 We're looking for more reviewers for the workshop! 📆 Review period: May 24-June 7 If you're passionate about making interpretability useful and want to help shape the conversation, we'd love your input. 💡🔍 Self-nominate here: docs.google.com/forms/d/e/1F...
Checkout our new preprint/project which has been over a year in the making! This has been a very fun collaboration (and one of the biggest I've personally participated in). @amuuueller.bsky.social @boknilev.bsky.social and other co-authors are around #ICLR2025 if you want to find out more. 😊
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
I'm not at #ICLR2025, but have 2 works being presented: 1) Understanding how LMs answer multiple-choice questions - arxiv.org/abs/2407.15018 - @boknilev.bsky.social is presenting the poster *now* until 12:30 (Hall 3+Hall 2B #207) - & w/ @oyvind-t.bsky.social @hanna-nlp.bsky.social Ashish Sabharwal
I'm in Singapore for ICLR to present this paper: Tomorrow, April 26th, 10-12:30 in Hall 3+2B #236 Come check it out! arxiv.org/abs/2504.12459
💡 New ICLR paper! 💡 "On Linear Representations and Pretraining Data Frequency in Language Models": We provide an explanation for when & why linear representations form in large (or small) language models. Led by @jackmerullo.bsky.social, w/ @nlpnoah.bsky.social & @sarah-nlp.bsky.social
Have work on the actionable impact of interpretability findings? Consider submitting to our Actionable Interpretability workshop at ICML! See below for more info. Website: actionable-interpretability.github.io Deadline: May 9
🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!
📢 Open PhD Position in Interpretable Natural Language Processing at the Department of Computer Science, UCPH! 🗓 Application deadline is 15 January 2025. Find more information about the position and apply here 👉 di.ku.dk/english/abou... @apepa.bsky.social @iaugenstein.bsky.social