SAEs give us fine-grained control over LLMs. How can we permanently encode feature ablations into an LM's parameters? We propose CRISP, and show that this improves unlearning over the prior state-of-the-art. Chat with @tomerashuach.bsky.social at ACL!
@tomerashuach.bsky.social
Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵
🚨New paper at #ACL2025 Findings! REVS: Unlearning Sensitive Information in LMs via Rank Editing in the Vocabulary Space. LMs memorize and leak sensitive data—emails, SSNs, URLs from their training. We propose a surgical method to unlearn it. 🧵👇w/ @boknilev.bsky.social @mtutek.bsky.social 1/8