Andreas Waldis

@tresiwald.bsky.social

Behavioral and Internal Interpretability 🔎 PostDoc Tübingen University | previously PhD Student at @ukplab.bsky.social, TU Darmstadt/Hochschule Luzern

I already presented some work on reference (names, pronouns, coreference resolution, pronoun fidelity, etc.) as a rich site to evaluate biases and commonsense reasoning, and our work on disentangling model behaviour and internals through aligned probing (led by @tresiwald.bsky.social).

Excited to present this work together with @dippedrusk.com at #EACL. Join us in the poster session 1 (11:30-13:00) 🔥

Poster of the Paper Aligned Probing
Andreas Waldis@tresiwald.bsky.social · 6mo ago

LMs that "know more" about toxicity are less toxic! Our #TACL 📄 connects behavior and internals: 💠 LMs amplify toxicity beyond humans 💠 Information about toxicity peaks in lower layers 💠 Bypassing these layers increases toxicity More details👇 #NLProc #interpretability (1/🧵)

simplified overview of our aligned probing setup, where we join the behavioral and internal evaluation of LMs' toxicity

Thanks a lot to everyone for the support, guidance, mentoring, collaboration, and great moments over the past years! 🙏 Without you, this journey wouldn't have been such a pleasure — and now excited to see what the future brings! 🚀

UKP Lab@ukplab.bsky.social · 5mo ago

🎓 𝗣𝗵𝗗 𝗱𝗲𝗳𝗲𝗻𝘀𝗲 𝗰𝗼𝗻𝗴𝗿𝗮𝘁𝘂𝗹𝗮𝘁𝗶𝗼𝗻𝘀 𝘁𝗼 𝗔𝗻𝗱𝗿𝗲𝗮𝘀 𝗪𝗮𝗹𝗱𝗶𝘀! Congratulations to @tresiwald.bsky.social on the successful defense of his dissertation “𝘈𝘯 𝘐𝘯𝘵𝘦𝘨𝘳𝘢𝘭 𝘝𝘪𝘦𝘸 𝘰𝘯 𝘵𝘩𝘦 𝘙𝘦𝘭𝘪𝘢𝘣𝘪𝘭𝘪𝘵𝘺 𝘰𝘧 𝘓𝘢𝘯𝘨𝘶𝘢𝘨𝘦 𝘔𝘰𝘥𝘦𝘭𝘴 𝘧𝘰𝘳 𝘊𝘰𝘮𝘱𝘶𝘵𝘢𝘵𝘪𝘰𝘯𝘢𝘭 𝘈𝘳𝘨𝘶𝘮𝘦𝘯𝘵𝘢𝘵𝘪𝘰𝘯”, held at Technische Universität Darmstadt on February 17th, 2026.