Sebastian Bordt

@sbordt.bsky.social

Language models and interpretable machine learning. Postdoc @ Uni Tübingen. https://sbordt.github.io/

Our spotlight paper is happening today at the #NeurIPS poster session! Drop by if you want to chat about the nitty-gritty details of large-scale transformer training!

Leena C Vankadara@leenacvankadara.bsky.social · 8mo ago

📄 Paper: arxiv.org/abs/2505.22491 Catch our Spotlight at #NeurIPS2025 Today! 📅 Wed Dec 3 🕟 4:30 - 7:30 PM 📍 Exhibit Hall C,D,E — Poster #3903 Huge thanks to my amazing collaborators: @mohaas.bsky.social @sbordt.bsky.social @ulrikeluxburg.bsky.social

T

We need new rules for publishing AI-generated research. The teams developing automated AI scientists have customarily submitted their papers to standard refereed venues (journals and conferences) and to arXiv. Often, acceptance has been treated as the dependent variable. 1/

Our #ICML position paper: #XAI is similar to applied statistics: it uses summary statistics in an attempt to answer real world questions. But authors need to state how concretely (!) their XAI statistics contributes to answer which concrete (!) question! arxiv.org/abs/2402.02870

Sebastian Bordt@sbordt.bsky.social · last yr.

During the last couple of years, we have read a lot of papers on explainability and often felt that something was fundamentally missing🤔 This led us to write a position paper (accepted at #ICML2025) that attempts to identify the problem and to propose a solution. arxiv.org/abs/2402.02870 👇🧵

In explainable machine learning, we mostly have negative results for what post-hoc explanations cannot do. This work presents a surprisingly strong positive result for SHAP, showing that a simple sampling modification allows to reliably detect features that don't influence the model.

Ulrike Luxburg@ulrikeluxburg.bsky.social · last yr.

Ever aggregated SHAP values across sample points? Our #COLT2025 paper proves that this might be safe when your goal is to discard unimportant features - but only if you add one extra line of code that reshuffles your data! With Robi Bhattacharjee and Karolin Frohnapfel arxiv.org/abs/2503.23111

Is the distinction between "aleatoric" and "epistemic" uncertainty practically meaningful (or even well defined) in any real sense? Aleotoric uncertainty refers to irreducible unpredictability (e.g. unrealized randomness in nature) whereas epistemic refers to model uncertainty.

I really like coding with LLMs. This week Claude & ChatGPT convinced me my code was too slow. After 2 days of investigation, I think my code is just fine. Never again will I blindly trust you with my profiler logs!🤖

I just asked aistudio.google.com to write a review for a paper that we will submit to ICML. It's impressive. I believe with this tool, I could produce a mediocre paper review for almost any paper in less than 10 minutes (judged by the standard of reviews that we have at ML conferences).

ICML 2025 has some exciting changes. Here are two of my favorites. 1. Only 1 round of back-and-forth between authors & reviewers. The review process should not be an endless back and forth. It shouldn't be possible to get your paper accepted by exhausting reviewers.

Bild

Are you interested in data contamination and LLM benchmarks?🤖 Check out our poster today at the NeurIPS ATTRIB workshop (3-4:30pm)! 💡 TL;DR: In the large-data regime, a few times of data contamination matter less than you might think.

BildBild