🚩 AI-assisted scientific writing faces increasing scrutiny Can human or LLM editing fix it? 👇 240k edits on scientific abstracts expose LLM writing flaws, and show that human editing fails to fix them cc Sanchaita Hazra, Doeun Lee, @shocheen.bsky.social, Bodhisattwa Prasad Majumder 🧵
@patqdasilva.bsky.social
Super grateful to have received senior area chair highlight at #ACL2025NLP ⏳ The generalization of interpretability-based steering methods is at an inflection point 🚂 As a community, we need to place stronger emphasis on evaluating the reliability of methods if we care about long-term impact
🌟Excited to announce that “Steering off Course” was accepted to #ACL2025NLP for an Oral and Panel Discussion! arxiv.org/abs/2504.04635 📍Wed, 9AM, Level 2 Hall A 🍁I will also share this work at Actionable Interpretability @ActInterp at #ICML2025 📍Sat, 1PM, East Ballroom A
Steering language models by directly intervening on internal activations is appealing–but does it generalize? We study 3 popular steering methods with 36 models from 14 families (1.5-70B), exposing brittle performance and fundamental flaws in underlying assumptions 🧵👇 (1/10)
Steering language models by directly intervening on internal activations is appealing–but does it generalize? We study 3 popular steering methods with 36 models from 14 families (1.5-70B), exposing brittle performance and fundamental flaws in underlying assumptions 🧵👇 (1/10)