Xinpeng Wang

@xinpeng.bsky.social

PhD student @LMU. Eval & LLM Alignment. https://xinpeng-wang.github.io/

Upcoming ICLR 2025 paper: ✂️ Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation We propose a surgical & flexible approach to mitigate false refusal in LLMs with minimal effect on performance and inference cost led by @xinpeng.bsky.social (1/2)