In our new paper, "A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages", we go beyond final-answer accuracy to analyze multilingual reasoning along three dimensions: performance, consistency, and faithfulness.
📝 What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns 👥 @mhedderich.bsky.social Anyi Wang @raoyuan.bsky.social @florian-eichin.com Jonas Fischer @barbaraplank.bsky.social 🔗 arxiv.org/abs/2504.158... 📁Main - Long
What changes if you take the LLM prompt “Tell me a short story about Dr. Li” and replace “Dr. Li” with “Dr. Smith”? Would you have guessed that this introduces a massive gender bias, from ca. half/half to 99% male doctors? In our #ACL2025 paper we present the Spotlight framework which...
Caught some great moments at #MCML Munich AI Day 2025 last week📍 From sharp keynotes to poster debates. Our team had the chance to show some recent work, join the conversations, and bring back plenty of food for thought🧠🗣️📊
Last week, #MCML Munich AI Day 2025 kicked off with keynotes by Julia Schnabel and Tina Eliassi-Rad, brilliantly moderated by Eva Schulz.
Dei Boarisch heard ned bei "Servus" und "Pfiade" auf? Dann suach ma genau Di! Wir suachan Bairischsprecher:innen, de a kurze Umfrage über KI-generierds Boarisch für a Masterarbeit beantwortn mechadn. Mid jeder Teilnahme bring ma den boarischn Dialekt a Stickal weida in de digitale Weyd!
Bavarian dialect speakers needed! Our MSc student Miriam wants to find out 1. how good/bad LLM-generated "Bavarian" is, and 2. whether dialect speakers agree with each other on this. The survey takes <5 min: survey.ifkw.lmu.de/dialquali25/ Thank you for sharing/participating!
Want to know if your prompting is also affected by this? Addressing this and other issues systematically, we proposed Spotlight, which utilizes data mining to uncover the effects of prompt- and model-changes (meet us at ACL to discuss) arxiv.org/abs/2504.15815
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts. Existing evaluation methods, eithe...
arxiv.org
a question mark changes the response. truly incredible