Marius Mosbach

@mariusmosbach.bsky.social

#NLP Postdoc at Mila - Quebec AI Institute & McGill University mariusmosbach.com

Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵

Paper title "The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?" with the paper's graphical abstract showing how more powerful alignment maps between a DNN and an algorithm allow more complex features to be found and more "accurate" abstractions.

💡Beyond math/code, instruction following with verifiable constraints is suitable to be learned with RLVR. But the set of constraints and verifier functions is limited and most models overfit on IFEval. We introduce IFBench to measure model generalization to unseen constraints.

Bild

A blizzard is raging through Montreal when your friend says “Looks like Florida out there!” Humans easily interpret irony, while LLMs struggle with it. We propose a 𝘳𝘩𝘦𝘵𝘰𝘳𝘪𝘤𝘢𝘭-𝘴𝘵𝘳𝘢𝘵𝘦𝘨𝘺-𝘢𝘸𝘢𝘳𝘦 probabilistic framework as a solution. Paper: arxiv.org/abs/2506.09301 to appear @ #ACL2025 (Main)

Bild

Started a new podcast with @tomvergara.bsky.social ! Behind the Research of AI: We look behind the scenes, beyond the polished papers 🧐🧪 If this sounds fun, check out our first "official" episode with the awesome Gauthier Gidel from @mila-quebec.bsky.social : open.spotify.com/episode/7oTc...

02 | Gauthier Gidel: Bridging Theory and Deep Learning, Vibes at Mila, and the Effects of AI on Art

Behind the Research of AI · Episode

open.spotify.com

Interested in shaping the progress of responsible AI and meeting leading researchers in the field? SoLaR@COLM 2025 is looking for paper submissions and reviewers! 🤖 ML track: algorithms, math, computation 📚 Socio-technical track: policy, ethics, human participant research

Bild

Excited to share the results of my recent internship! We ask 🤔 What subtle shortcuts are VideoLLMs taking on spatio-temporal questions? And how can we instead curate shortcut-robust examples at a large-scale? We release: MVPBench Details 👇🔬

Bild

Chain-of-Thought (CoT) reasoning lets LLMs solve complex tasks, but long CoTs are expensive. How short can they be while still working? Our new ICML paper tackles this foundational question.

Bild

I'll be at #NAACL2025: 🖇️To present my paper "Superlatives in Context", showing how the interpretation of superlatives is very context dependent and often implicit, and how LLMs handle such semantic underspecification 🖇️And we will present RewardBench on Friday Reach out if you want to chat!

Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!

Logo for MIB: A Mechanistic Interpretability Benchmark

Checkout Benno's notes about our impact of interpretability paper 👇. Also, we are organizing a workshop at #ICML2025 which is inspired by some of the questions discussed in the paper: actionable-interpretability.github.io

General Information

ICML 2025 - Vancouver

actionable-interpretability.github.io

Benno Krojer@bennokrojer.bsky.social · last yr.

Day 12: From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP arxiv.org/abs/2406.12618 Genuinely one of my favourite papers in recent years! It tries to answer one question that every phd student often asks themselves: Does this research matter?

Introducing nanoAhaMoment: Karpathy-style, single file RL for LLM library (<700 lines) - super hackable - no TRL / Verl, no abstraction💆‍♂️ - Single GPU, full param tuning, 3B LLM - Efficient (R1-zero countdown < 10h) comes with a from-scratch, fully spelled out YT video [1/n]

Bild

Check out our new workshop on Actionable Interpretability @ ICML 2025. We are also looking forward to submissions that take a position on the future of interpretability research more broadly. 👇

Mor Geva@megamor2.bsky.social · last yr.

🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!