🚨 New paper alert! WaX, a method to explain what drives dataset shift by decomposing Wasserstein distances into exact instance- and feature-level attributions using XAI. doi.org/10.1109/TPAM...
Explainable AI Berlin
@xai-berlin.bsky.social
Explainable AI research from the machine learning group of Prof. Klaus-Robert Müller at @tuberlin.bsky.social & @bifold.berlin
🚨 New preprint alert! Distributed Sparse Interventions (DSI), a method to steer LLM behavior by intervening on as few as 8-64 neurons (as little as 0.01% of a model!), instead of whole layers or directions in activation space. 🔗 arxiv.org/abs/2607.07128
🔦 Past Paper Highlight: PRISM generates multi-concept descriptions of LLM features by clustering high-activation inputs and labeling each cluster. Two evaluation metrics measure the semantic diversity and faithfulness of the resulting descriptions. 🔗 proceedings.neurips.cc/paper_files/...
📣 Announcing the BlackboxNLP 2026 Reproducibility Challenge! A new track dedicated to rigorous robustness checks of NLP interpretability work - stress-testing baselines, ablations, generalizability, and evaluation.
🚨 New Paper Alert: Hidden inside every SVM and k-nearest-neighbors classifier is a neural-network structure. Exposing it makes fast, faithful LRP explanations available for distance-based classifiers, where gradient methods fail. 📄 doi.org/10.1016/j.pa...
🔦Past Paper Highlight: Scaling Higher-Order Explanations for Graph Neural Networks 📄Efficient Higher-Order Subgraph Attribution via Message Passing 🔗 proceedings.mlr.press/v162/xiong22... 📄Relevant Walk Search for Explaining Graph Neural Networks 🔗 proceedings.mlr.press/v202/xiong23...
🔦 Past Paper Highlight: GNN-LRP, explaining Graph Neural Networks via higher-order interactions 📄 Higher-Order Explanations of Graph Neural Networks via Relevant Walks 🔗 ieeexplore.ieee.org/document/954... 💻 git.tu-berlin.de/thomas_schna...
🔦 Past Paper Highlight: PRCA & DRSA, methods that disentangle neural network explanations into concept-level components. 📄 Disentangled Explanations of Neural Network Predictions by Finding Relevant Subspaces 🔗 doi.org/10.1109/TPAM... 💻 github.com/p16i/disenta...
BIFOLD is hiring — 3 open positions in machine learning & security, data management, and biomedical sensing. 🧵 t1p.de/ikhv5 #phdsky #fundedphd #neurojobs #sciencejobs #STEMjobs #academicjobs @tuberlin.bsky.social #AcademicSky #STEMJobs, #ScienceJobs, #SciJobs #NeuroJobs #AcademicJobs #fNirs 📈🤖
BIFOLD Research Job Opportunities / Join us in Berlin
BIFOLD welcomes job applications year-round Ideal candidates are those who seek to challenge themselves academically, are passionate about data science and computer systems engineering, and possess ...
t1p.de
🚨 Final call: 10 PhD positions in ML & Data Science at BIFOLD Application deadline: February 13, 2025 – This Friday! #GraduateSchool 2026 www.jobs.tu-berlin.de/en/job-posti... #hiring #phd #AcademicJobs #PhDPosition #naturalscience #AI #phd #phdjobs #VacancyEdu #ScienceCareer
#graduateschool #datamanagement #machinelearning #phd #ai #machinelearning #datascience #research #berlin #bifold #academiccareers #doctoralresearch | BIFOLD - Berlin Institute for the Foundations of ...
🚨 Final call: 10 PhD positions in AI & Data Science at BIFOLD Berlin Application deadline: February 13, 2025 – This Friday! #GraduateSchool 2026 https://lnkd.in/duheJE8J The Berlin Institute for th...
linkedin.com
🔦 Past Paper Highlight: BiLRP, explaining Machine Learning Models Through Feature Interactions 📄 Building and Interpreting Deep Similarity Models 🔗 ieeexplore.ieee.org/document/918... 💻 github.com/oeberle/BiLR...
Last week, Dr. Nils Feldhus @nfel.bsky.social, postdoctoral researcher at @tuberlin.bsky.social and @bifold.berlin, visited our lab and presented his research during our weekly lab meeting.
🔎 Today, we highlight a brand new paper, introducing a triad of open-source packages for dataset-wide and quantitative analyses of attributions. Attribution methods are powerful explanation tools, but by themselves cumbersome to generate dataset-wide insights. Not anymore!
2025 marks the 10-year anniversary of Layer-wise Relevance Propagation (LRP)! 🥳 To celebrate a decade of this attribution method, we highlight key milestones from the past ten years in our LinkedIn post. Learn more here: www.linkedin.com/posts/xai-be...
Klaus-Robert Müller, Pioneer in Machine Learning, honored with the 2026 Gottfried Wilhelm Leibniz Prize by @dfg.de www.bifold.berlin/news-events/... @tuberlin.bsky.social #mlsky @xai-berlin.bsky.social #datascience #AI
San Diego 🇺🇸 or Mexico City 🇲🇽 for #NeurIPS2025? We got you covered either way 😎 On Dec 3rd: 🇲🇽 @dilya.bsky.social present our work on the fragility of Mech Interp in Mexico 🇺🇸 @lkopf.bsky.social present our work on polysemanticity in San Diego I am not there this year, so I‘ll be cheering from afar!
✈️🇲🇽 Next Wednesday (Dec 3), 1–4 p.m. CST, I’ll be presenting Manipulating Feature Visualizations with Gradient Slingshots at NeurIPS 2025 in Mexico City! Feature Visualization has long been a staple interpretability tool. Our work shows it’s far from reliable! 🚨
Join us today from 4:30 to 7:30 PM @neuripsconf.bsky.social Hall C,D,E #1006 for our poster on SmoothDiff, a novel XAI method leveraging automatic differentiation. 🧵1/6
I’m at #NeurIPS in San Diego this week! Come see our poster on feature interpretability. Find @eberleoliver.bsky.social and me at: 🪧Poster Session 1 @ Exhibit Hall C,D,E #1015 Wed 3 Dec, 11 am - 2 pm 🪧Poster @ Mech Interp Workshop Upper Level Room 30A-E Sun 7 Dec, 8 am - 5 pm
🚀 Visit our #NeurIPS posters at @neuripsconf.bsky.social! Meet and interact with our authors at all locations — San Diego, Mexico City, and Copenhagen. Details in the thread. 👇👇👇
Congratulations to Jonas Dippel on completing his PhD. 🎉 Alongside his work on histopathology foundation models 🧱, he contributed to multiple projects in explainable AI 🔎, advancing precision pathology and uncovering spurious correlations in unsupervised vision models. #XAI #MachineLearning
Heading to the EMNLP BlackboxNLP Workshop this Sunday? Don’t miss @nfel.bsky.social and @lkopf.bsky.social poster on „Interpreting Language Models Through Concept Descriptions: A Survey“ aclanthology.org/2025.blackbo... #EMNLP #BlackboxNLP #XAI #Interpretapility
Nov 9, @blackboxnlp.bsky.social , 11:00-12:00 @ Hall C – Interpreting Language Models Through Concept Descriptions: A Survey (Feldhus & Kopf) @lkopf.bsky.social 🗞️ aclanthology.org/2025.blackbo... bsky.app/profile/nfel...
We are grateful for the opportunity to present some of our work at the All Hands Meeting of the German AI Centers, hosted by @dfki.bsky.social in Saarbrücken. Andreas Lutz @eberleoliver.bsky.social Manuel Welte @lorenzlinhardt.bsky.social @lkopf.bsky.social #AI #XAI #Interpretability
Happy to share that our PRISM paper has been accepted at #NeurIPS2025 🎉 In this work, we introduce a multi-concept feature description framework that can identify and score polysemantic features. 📄 Paper: arxiv.org/abs/2506.15538 #NeurIPS #MechInterp #XAI
This is the eXplainable AI research channel of the machine learning group of Prof. Klaus-Robert Müller at Technische Universität Berlin @tuberlin.bsky.social & BIFOLD @bifold.berlin. Let's connect! #XAI #ExplainableAI #MechInterp #MachineLearning #Interpretability
a black background with green text that says `` hello , world ''
ALT: a black background with green text that says `` hello , world ''
media.tenor.com