Gabriele Sarti

@gsarti.com

Open-source interpretability to seize the means of prediction. Postdoc @ Northeastern, @ndif-team.bsky.social w/ @davidbau.bsky.social. gsarti.com

⏳ The BlackboxNLP 2026 Reproducibility Challenge deadline has been extended to July 24 (AoE) ⏳ If you've been working on a robustness check, ablation, or replication of recent NLP interpretability work, you have a bit more time to get your submission in.

Bild

I'd never have guessed models commit to their final answer this early, often within the first 20% of reasoning, across math/logic tasks and model families. The rest is mostly hedging that doesn't change their mind. And turns out they encode this internally, we can decode it! 🧵👇

Sara Candussio@saracandussio.bsky.social · last mo.

Are all the CoT steps necessary? In our latest paper, we find evidence for the existence of a commitment boundary, marking a sharp transition from no/mid guesses to the model final answer across various reasoning tasks and model families. Thread 🧵👇

🚨New paper led by @danielsc4.it & @saracandussio.bsky.social studying answer commitment across CoT for math & logic reasoning! Turns out, models often converge to final answers early in the CoT, then tend to "fake" hedging & re-checks - and we can tell apart mid/final guesses quite robustly! See 👇

Sara Candussio@saracandussio.bsky.social · last mo.

Are all the CoT steps necessary? In our latest paper, we find evidence for the existence of a commitment boundary, marking a sharp transition from no/mid guesses to the model final answer across various reasoning tasks and model families. Thread 🧵👇

Our @ndif-team.bsky.social was thrilled to sponsor the BlackboxNLP repro challenge! Interpretability needs more robust & generalizable findings - if you're working in this space, this should be on your radar!

BlackboxNLP@blackboxnlp.bsky.social · last mo.

🏆 Announcing the NDIF Best Paper Award for the BlackboxNLP 2026 Reproducibility Challenge! @ndif-team.bsky.social Reproduce an interp. finding with nnsight + NDIF, open-source it, and push it further. Winners will receive a $500 prize, and will be invited to present at the workshop!

"the new toaster says 'I Love You' when you put the toast in, and 'I'm Sorry' when it burns it. This causes some people to get angry b/c they don't think the toaster means it, and others to develop unhealthy attachments b/c they think it does. the solution is to make the toaster TRULY sorry"

The New England Mechanistic Interpretability (NEMI) workshop is coming to BU on Aug. 14! Join us for talks, a panel, food, and plenty of opportunities to connect with the many great researchers in the area. Register and help spread the word!

Micah Benson@micahben.bsky.social · 2mo ago

🧠🤖 The 2026 New England Mechanistic Interpretability (NEMI) Workshop will be Aug. 14 at Boston University! Help spread the word and join the New England mech interp community! Registration and submission info in thread:👇

BU campus and Boston skyline

New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...

I am worried by NLP research culture

In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…

ehudreiter.com

D

"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.

Bild

With the large influx of submissions and a faster pace of research, reproducibility is more important than ever. With this reproducibility challenge, we want to put the focus on best practices wrt. baselines🧱, ablations🌈, eval🔎 and generalizability🗺️ of interpretability!

BlackboxNLP@blackboxnlp.bsky.social · 3mo ago

📣 Announcing the BlackboxNLP 2026 Reproducibility Challenge! A new track dedicated to rigorous robustness checks of NLP interpretability work - stress-testing baselines, ablations, generalizability, and evaluation.

Despite the huge inflow of researchers, much of the work in interpretability remains anecdotal. Our new repro challenge at BlackboxNLP (co-located with EMNLP 2026) aims to attract work challenging common assumptions and showing failure/success cases of popular methods. Negative results welcome!

BlackboxNLP@blackboxnlp.bsky.social · 3mo ago

📣 Announcing the BlackboxNLP 2026 Reproducibility Challenge! A new track dedicated to rigorous robustness checks of NLP interpretability work - stress-testing baselines, ablations, generalizability, and evaluation.

IMO a key skill of a good scientist is moving comfortably between the specific & general e.g., relentless in understand the nerdy details of the data (including its coding) WHILE holding onto the big picture of the hypotheses being tested with the data

Today we’re releasing EMO, a new mixture-of-experts (MoE) model trained so modular structure emerges directly from data without human-defined priors. EMO can use a small subset of its experts for a given task while keeping near full-model performance. 🧵

Bild

New paper w/ UK AISI: Millions of people now use AI to help them write and communicate. In three experiments (14k participants, 3m+ human ratings) we show that AI writing assistance systematically distorts writer personas – their perceived beliefs, personality, and identity. 🧵

Bild