Yonatan Belinkov ✈️ COLM2025

@boknilev.bsky.social

Associate professor of computer science at Technion; visiting scholar at @KempnerInst 2025-2026 https://belinkov.com/

📣 Announcing the BlackboxNLP 2026 Reproducibility Challenge! A new track dedicated to rigorous robustness checks of NLP interpretability work - stress-testing baselines, ablations, generalizability, and evaluation.

Bild

BlackboxNLP will be co-located with EMNLP 2026 in 🇭🇺 Budapest 🇭🇺 this October! This edition will feature a special reproducibility track, investigating generalization and robustness of established results from interpretability research 👷‍♂️ Stay tuned for more details!

Bild

🤔What happens when LLM agents choose between achieving their goals and avoiding harm to humans in realistic management scenarios? Are LLMs pragmatic or prefer to avoid human harm? 🚀 New paper out: ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs🚀🧵

Bild

What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics!

Bild

Excited to join @KempnerInst this year! Get in touch if you're in the Boston area and want to chat about anything related to AI interpretability, robustness, interventions, safety, multi-modality, protein/DNA LMs, new architectures, multi-agent communication, or anything else you're excited about!

Kempner Institute at Harvard University@kempnerinstitute.bsky.social · 11mo ago

News from the #KempnerInstitute! We’re thrilled to welcome Yonatan Belinkov (expert in #NLP) and Daphna Weinshall (expert in human & machine vision) as visiting scholars for the 2025–26 academic year. 📖 Read more: bit.ly/47QkDID #AI #MachineVision @boknilev.bsky.social

BlackboxNLP is the workshop on interpreting and analyzing NLP models (including LLMs, VLMs, etc). We accept full (archival) papers and extended abstracts. The workshop is highly attended and is a great exposure for your finished work or feedback on work in progress. #emnlp2025 at Sujhou, China!

BlackboxNLP@blackboxnlp.bsky.social · 12mo ago

📢 Call for Papers! 📢 #BlackboxNLP 2025 invites the submission of archival and non-archival papers on interpreting and explaining NLP models. 📅 Deadlines: Aug 15 (direct submissions), Sept 5 (ARR commitment) 🔗 More details: blackboxnlp.github.io/2025/call/

Join our Discord for discussions and a bunch of simple submission ideas you can try! discord.gg/n5uwjQcxPR Participants will have the option to write a system description paper that gets published.

Join the BlackboxNLP Shared Task Discord Server!

Check out the BlackboxNLP Shared Task community on Discord - hang out with 83 other members and enjoy free voice and text chat.

discord.gg

BlackboxNLP@blackboxnlp.bsky.social · last yr.

⏳ Three weeks left! Submit your work to the MIB Shared Task at #BlackboxNLP, co-located with @emnlpmeeting.bsky.social Whether you're working on circuit discovery or causal variable localization, this is your chance to benchmark your method in a rigorous setup!

Have you started working on your submission for the MIB shared task yet? Tell us what you’re exploring! New featurization methods? Circuit pruning? Better feature attribution? We'd love to hear about it 👇

Working on feature attribution, circuit discovery, feature alignment, or sparse coding? Consider submitting your work to the MIB Shared Task, part of this year’s #BlackboxNLP We welcome submissions of both existing methods and new or experimental POCs!

Bild

VLMs perform better on questions about text than when answering the same questions about images - but why? and how can we fix it? In a new project led by Yaniv (@YNikankin on the other app), we investigate this gap from an mechanistic perspective, and use our findings to close a third of it! 🧵

Bild

Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵

Bild

1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇

Bild