Andrew Lee

@ajyl.bsky.social

Post-doc @ Harvard. PhD UMich. Spent time at FAIR and MSR. ML/NLP/Interpretability

If you liked Anthropic's recent emotions paper, check out our work! We find many similarities: 1) Circular geometry of emotion representations 2) Steering: unlike Anthropic, we steer along circular manifold (at 0°, 30°, 60°...) 3) Steering emotions can affect refusal/sycophancy See Lihao's thread!👇

Lihao Sun@1e0sun.bsky.social · 4mo ago

💡New paper! Woke up to Anthropic's emotion paper and realized “wait, that's our finding too.” We concurrently uncovered a circular valence & arousal (VA) geometry of emotions, steering refusal & sycophancy. We further provide a mechanistic account: tokens occupy distinct regions in this space. 1/

We are thrilled to host the next Mech Interp Workshop @ ICML 2026! 🎉 July 2026, Seoul 🇰🇷 The workshop aims to understand the inner workings of neural nets. Topics: Feature geometry Circuit analyses Interp for {practical applications, safety, scientific discovery}, and many more.

Bild

Question @neuripsconf.bsky.social - a coauthor had his reviews re-assigned many weeks ago. The ACs of those papers told him "i've been told to tell u: leave a short note. You won't be penalized". Now I'm being warned of desk-reject due to his short/poor reviews. What's the right protocol here?

How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!

Bild

🚨New #ACL2025 paper! Today’s “safe” language models can look unbiased—but alignment can actually make them more biased implicitly by reducing their sensitivity to race-related associations. 🧵Find out more below!

Bild

🚨New preprint! How do reasoning models verify their own CoT? We reverse-engineer LMs and find critical components and subspaces needed for self-verification! 1/n

Bild

🚨New Preprint! Did you know that steering vectors from one LM can be transferred and re-used in another LM? We argue this is because token embeddings across LMs share many “global” and “local” geometric similarities!

Bild

New paper <3 Interested in inference-time scaling? In-context Learning? Mech Interp? LMs can solve novel in-context tasks, with sufficient examples (longer contexts). Why? Bc they dynamically form *in-context representations*! 1/N

Bild