Shauli Ravfogel

@shauli.bsky.social

Faculty fellow at NYU CDS. Previously: PhD @ BIU NLP.

1/ Can LLMs introspect, i.e., reason about their internal states? Recent work claims LLMs notice when their "thoughts" get tampered with, and can report the content. We took a closer look and think it's too early to say that. Work led by Shashwat Singh, with @tallinzen.bsky.social and me. A thread 🧵

Bild

I’ll be at NeurIPS in San Diego presenting this paper during the Wednesday, Dec 3 poster session (11 am – 2 pm PST) & at the mechanistic interpretability workshop on Sunday (spotlight). Come say hi, and feel free to DM if you’d like to talk research or just catch up!

Shauli Ravfogel@shauli.bsky.social · 10mo ago

New NeurIPS paper! Why do LMs represent concepts linearly? We focus on LMs's tendency to linearly separate true and false assertions, and provide an analysis of the truth circuit in a toy model. A joint work with Gilad Yehudai, @tallinzen.bsky.social, Joan Bruna and @albertobietti.bsky.social.

How well can LLMs understand tasks with complex sets of instructions? We investigate through the lens of RELIC: REcognizing (formal) Languages In-Context, finding a significant overhang between what LLMs are able to do theoretically and how well they put this into practice.

Bild

A quick update: I’ve completed my PhD at Bar-Ilan University. After an amazing research visit in Prof. Ryan Cotterell’s lab at ETH Zurich, I am super excited to join NYU Center for Data Science as a Faculty Fellow!

Bild