1/ Can LLMs introspect, i.e., reason about their internal states? Recent work claims LLMs notice when their "thoughts" get tampered with, and can report the content. We took a closer look and think it's too early to say that. Work led by Shashwat Singh, with @tallinzen.bsky.social and me. A thread 🧵
Shauli Ravfogel
@shauli.bsky.social
Faculty fellow at NYU CDS. Previously: PhD @ BIU NLP.
LLM introspection revisited. if we do the controls properly we may not have strong enough evidence just yet arxiv.org/html/2605.26...
Can LLMs Introspect? A Reality Check
arxiv.org
I’ll be at NeurIPS in San Diego presenting this paper during the Wednesday, Dec 3 poster session (11 am – 2 pm PST) & at the mechanistic interpretability workshop on Sunday (spotlight). Come say hi, and feel free to DM if you’d like to talk research or just catch up!
New NeurIPS paper! Why do LMs represent concepts linearly? We focus on LMs's tendency to linearly separate true and false assertions, and provide an analysis of the truth circuit in a toy model. A joint work with Gilad Yehudai, @tallinzen.bsky.social, Joan Bruna and @albertobietti.bsky.social.
New NeurIPS paper! Why do LMs represent concepts linearly? We focus on LMs's tendency to linearly separate true and false assertions, and provide an analysis of the truth circuit in a toy model. A joint work with Gilad Yehudai, @tallinzen.bsky.social, Joan Bruna and @albertobietti.bsky.social.
1/8 Happy to share our new paper—“IQ Test for LLMs”—co-authored with Aviya Maimon, Amir DN Cohen, @neurogal.bsky.social and Reut Tsarfaty. We propose to rethink how language models are evaluated by focusing on the latent capabilities that explain benchmark results. Arxiv: arxiv.org/pdf/2507.20208
I’ll be at #ACL2025! If you’re around and want to catch up or chat, please ping me!
How well can LLMs understand tasks with complex sets of instructions? We investigate through the lens of RELIC: REcognizing (formal) Languages In-Context, finding a significant overhang between what LLMs are able to do theoretically and how well they put this into practice.
Introduction \ updates post (better late than never): - I recently graduated from ELSC @hebrewuniversity.bsky.social 🎉 - Moved to NYC 🗽 (view from our balcony below 👇) - And started a postdoc in @columbiauniversity.bsky.social Pls PM if you are in the NYC area and want to talk (or have a beer 🍻)
Our paper "A Practical Method for Generating String Counterfactuals" has been accepted to the findings of NAACL 2025! a joint work with @matan-avitan.bsky.social , @yoavgo.bsky.social and Ryan Cotterell. We propose "Intervention Lens", a technique to explain intervention in natural language. (1/6)
A quick update: I’ve completed my PhD at Bar-Ilan University. After an amazing research visit in Prof. Ryan Cotterell’s lab at ETH Zurich, I am super excited to join NYU Center for Data Science as a Faculty Fellow!
Happy to share our work "Counterfactual Generation from Language Models" with @AnejSvete, @vesteinns, and Ryan Cotterell! We tackle generating true counterfactual strings from LMs after interventions and introduce a simple algorithm for it. (1/7) arxiv.org/pdf/2411.07180