We have bigger models. Bigger conferences. Have our questions become bigger, too? That was the question I posed in my Presidential Address at #ACL2026 in San Diego. (1/6)
MaiNLP lab, LMU Munich
@mainlp.bsky.social
MaiNLP research lab at CIS, LMU Munich directed by Barbara Plank @barbaraplank.bsky.social Natural Language Processing | Artificial Intelligence | Computational Linguistics | Human-centric NLP
Super excited to present our work at #ACL2026 ✨ 👉 LLM-based social simulations have lots of analytic flexibility—response generation methods are often overlooked 📍Tue July 7, 9am, Poster Session F
Thrilled to share that our paper with @carohaensch.bsky.social, @barbaraplank.bsky.social & @mstrohm.bsky.social has been accepted to #ACL2026 Main! How to generate closed-ended survey responses 📊 with LLMs trained to produce open-ended text? 📝 We show results from 32 mio. simulated responses… 🧵
I'm at #ACL2026 where I'll present work on comparing text- vs. speech-based classification performance on standard-language vs. dialectal data, finding different trends for different varieties! (Talk @ Harbor A, July 5, 14:20-14:30) [1/3]
I’m at #ACL2026NLP to present ltzGLUE, a GLUE-style benchmark for Luxembourgish. We also release ltz-E1, a ModernBERT-based encoder trained only on Luxembourgish. Come find me on 7/7!
Excited for 5 papers at #ACL2026NLP with my group and with collaborators. 📍 You can find the work here: 🗓️ Sun. July 5 AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages Oral Session B: Multilinguality and Language Diversity 2
Want to learn more? Poster presentation at Tue, Jul 7, 2026, 2:00 PM – 3:45 PM @ HALL A #3303 Also in the workshop on mechanistic interpretability, Fri, Jul 10, 2026 • 8:00 AM – 5:00 PM @ HALL C Paper: arxiv.org/abs/2505.20076
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
Post-hoc interpretability methods typically attribute a model's behavior to its components, data, or training trajectory in isolation, and are often tied to a particular level of granularity along the...
arxiv.org
𝗚𝗿𝗮𝗱𝗶𝗲𝗻𝘁 𝗱𝗲𝘀𝗰𝗲𝗻𝘁 𝗶𝘀 𝗮𝗹𝗹 𝘆𝗼𝘂 𝗻𝗲𝗲𝗱 𝗳𝗼𝗿 𝗶𝗻𝘁𝗲𝗿𝗽𝗿𝗲𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆? In our ICML paper, we show that it might just be. ExPLAIND is a method that unifies data attribution, model component attribution, and training dynamics by computing an exact decomposition of model behavior based on gradient products.
MaiNLP will be at #ACL2026 with a number of talks & posters, plus the ACL presidential address! See you in San Diego 🌊
I'm happy to share that I successfully defended my PhD two days ago! Thank you very much to @barbaraplank.bsky.social & Hinrich Schütze for supervising me, to @dirkhovy.bsky.social for joining the committee, and to the many, many people who have been part of my PhD journey :)
🎉 Our paper “From 🔇 Noise to 📶 Signal to 🎯 Selbstzweck: Reframing Human Label” is accepted to Findings of ACL 2026! 👉 lnkd.in/eqDZkrBa ❤️ Huge thanks to my co-authors @santosh-tyss.bsky.social and @barbaraplank.bsky.social 🙏
MaiNLP will be at #LREC2026 to present multiple conference & workshop papers, as well as a workshop keynote! See you in Palma ☀️
MaiNLP is happy to be part of @eaclmeeting.bsky.social with several papers, talks, a panel, and a workshop ☀️ Looking forward to seeing you in Rabat! #EACL2026
We are honoured to welcome Prof Barbara Plank (@barbaraplank.bsky.social ) from @mainlp.bsky.social @cislmu.bsky.social, as our keynote speaker. LoResLM @eaclmeeting.bsky.social
📢 New paper accepted at @eaclmeeting.bsky.social 2026: Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions with @mhedderich.bsky.social @amodarressi.bsky.social Hinrich Schuetze & Benjamin Roth. Preprint: arxiv.org/abs/2512.12775
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
Expert persona prompting -- assigning roles such as expert in math to language models -- is widely used for task improvement. However, prior work shows mixed results on its effectiveness, and does not...
arxiv.org
In our new paper, "A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages", we go beyond final-answer accuracy to analyze multilingual reasoning along three dimensions: performance, consistency, and faithfulness.
✨New paper✨ We find script (e.g. Cyrillic, Latin) to be a linear direction in the activation space of Whisper, enabling transliteration at test-time by adding such script directions to the activations — producing e.g. Cyrillic Japanese transcriptions.
VarDial 2026 will be colocated with @eaclmeeting.bsky.social! We're looking forward to your papers on NLP for similar languages, varieties and dialects :) Deadline: Dec 19 (Jan 2 for pre-reviewed ARR papers) sites.google.com/view/vardial...
Congrats to Pingjun, @beiduo.bsky.social , Siyao, Marie, and @barbaraplank.bsky.social for receiving the SAC Highlights reward!
Senior Area Chair Highlights (8/9)
Congrats to our team member Diego Frassinelli on the SAC Highlights award!
Senior Area Chair Highlights (3/9)
At #Interspeech2025 I'm going to present Betthupferl, a dataset for German dialect ASR & dialect-to-standard speech translation! We analyze differences between dialectal & Standard German transcriptions, benchmark ASR models, and examine shortcomings of current ASR models & evaluation metrics.
UPDATE: Our poster presentation got moved to Tuesday, 16:00–17:30 (session 10)! #ACL2025NLP
At #ACL2025NLP I'll present our analysis of the effect of linguistic similarity on cross-lingual transfer! We looked at how 10 similarity measures correlate w/ transfer results btwn 263 languages across 3 NLP tasks. Different similarity measures matter for diff. experiments (no one-size-fits-all)!
Headed to ACL? MaiNLP & our most recent work will be there too👥📄 Come see what we’ve been working on!
📄 [ACL 2025 main] Circuit compositions: Exploring Modular Structures in Transformer-Based Language Models (doi.org/10.48550/arX...)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
A fundamental question in interpretability research is to what extent neural networks, particularly language models, implement reusable functions through subnetworks that can be composed to perform mo...
doi.org
📄 [ACL 2025 main] LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (doi.org/10.48550/arX...)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case...
doi.org
At #ACL2025NLP I'll present our analysis of the effect of linguistic similarity on cross-lingual transfer! We looked at how 10 similarity measures correlate w/ transfer results btwn 263 languages across 3 NLP tasks. Different similarity measures matter for diff. experiments (no one-size-fits-all)!
🤔 Can LLMs read between the lines? Our another #ACL2025 paper surveys resources on how LLMs handle pragmatics like implicatures, deixis, and more. We map out a new landscape for both LLMs and linguistics in pragmatic research. 📄 arxiv.org/abs/2502.12378 🧠💬 #LLMs #Pragmatics