🔦 Past Paper Highlight: PRISM generates multi-concept descriptions of LLM features by clustering high-activation inputs and labeling each cluster. Two evaluation metrics measure the semantic diversity and faithfulness of the resulting descriptions. 🔗 proceedings.neurips.cc/paper_files/...
Laura Kopf
@lkopf.bsky.social
PhD student in Interpretable Machine Learning at @tuberlin.bsky.social & @bifold.berlin https://web.ml.tu-berlin.de/author/laura-kopf/
Last week, Dr. Nils Feldhus @nfel.bsky.social, postdoctoral researcher at @tuberlin.bsky.social and @bifold.berlin, visited our lab and presented his research during our weekly lab meeting.
I’m at #NeurIPS in San Diego this week! Come see our poster on feature interpretability. Find @eberleoliver.bsky.social and me at: 🪧Poster Session 1 @ Exhibit Hall C,D,E #1015 Wed 3 Dec, 11 am - 2 pm 🪧Poster @ Mech Interp Workshop Upper Level Room 30A-E Sun 7 Dec, 8 am - 5 pm
🚀 Visit our #NeurIPS posters at @neuripsconf.bsky.social! Meet and interact with our authors at all locations — San Diego, Mexico City, and Copenhagen. Details in the thread. 👇👇👇
Heading to the EMNLP BlackboxNLP Workshop this Sunday? Don’t miss @nfel.bsky.social and @lkopf.bsky.social poster on „Interpreting Language Models Through Concept Descriptions: A Survey“ aclanthology.org/2025.blackbo... #EMNLP #BlackboxNLP #XAI #Interpretapility
Nov 9, @blackboxnlp.bsky.social , 11:00-12:00 @ Hall C – Interpreting Language Models Through Concept Descriptions: A Survey (Feldhus & Kopf) @lkopf.bsky.social 🗞️ aclanthology.org/2025.blackbo... bsky.app/profile/nfel...
We are grateful for the opportunity to present some of our work at the All Hands Meeting of the German AI Centers, hosted by @dfki.bsky.social in Saarbrücken. Andreas Lutz @eberleoliver.bsky.social Manuel Welte @lorenzlinhardt.bsky.social @lkopf.bsky.social #AI #XAI #Interpretability
Nov 9, @blackboxnlp.bsky.social , 11:00-12:00 @ Hall C – Interpreting Language Models Through Concept Descriptions: A Survey (Feldhus & Kopf) @lkopf.bsky.social 🗞️ aclanthology.org/2025.blackbo... bsky.app/profile/nfel...
Interpreting Language Models Through Concept Descriptions: A Survey
Nils Feldhus, Laura Kopf. Proceedings of the 8th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. 2025.
aclanthology.org
🔍 Are you curious about uncovering the underlying mechanisms and identifying the roles of model components (neurons, …) and abstractions (SAEs, …)? We provide the first survey of concept description generation and evaluation methods. Joint effort w/ @lkopf.bsky.social 📄 arxiv.org/abs/2510.01048
This is the eXplainable AI research channel of the machine learning group of Prof. Klaus-Robert Müller at Technische Universität Berlin @tuberlin.bsky.social & BIFOLD @bifold.berlin. Let's connect! #XAI #ExplainableAI #MechInterp #MachineLearning #Interpretability
a black background with green text that says `` hello , world ''
ALT: a black background with green text that says `` hello , world ''
media.tenor.com
🔍 Are you curious about uncovering the underlying mechanisms and identifying the roles of model components (neurons, …) and abstractions (SAEs, …)? We provide the first survey of concept description generation and evaluation methods. Joint effort w/ @lkopf.bsky.social 📄 arxiv.org/abs/2510.01048
Happy to share that our PRISM paper has been accepted at #NeurIPS2025 🎉 In this work, we introduce a multi-concept feature description framework that can identify and score polysemantic features. 📄 Paper: arxiv.org/abs/2506.15538 #NeurIPS #MechInterp #XAI
🔍 When do neurons encode multiple concepts? We introduce PRISM, a framework for extracting multi-concept feature descriptions to better understand polysemanticity. 📄 Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework arxiv.org/abs/2506.15538 🧵 (1/7)
Still overwhelmed by the amazing response to our poster session at @neuripsconf.bsky.social with Anna Hedström and Marina Höhne! It was incredible to have such lively and inspiring discussions with brilliant people whose work I admire. ✨
I’ll be presenting our work at @neuripsconf.bsky.social in Vancouver! 🎉 Join me this Thursday, December 12th, in East Exhibit Hall A-C, Poster #3107, from 11 a.m. PST to 2 p.m. PST. I'll be discussing our paper “CoSy: Evaluating Textual Explanations of Neurons.”