1/ 🌍 How does mixing data from hundreds of languages affect LLM training? In our new paper "Revisiting Multilingual Data Mixtures in Language Model Pretraining" we revisit core assumptions about multilinguality using 1.1B-3B models trained on up to 400 languages. 🧵👇
🚨New Preprint! In multilingual models, the same meaning can take far more tokens in some languages, penalizing users of underrepresented languages with worse performance and higher API costs. Our Parity-aware BPE algorithm is a step toward addressing this issue: 🧵
Super excited to share that our paper "A Logical Fallacy-Informed Framework for Argument Generation" has received the Outstanding Paper Award 🎉🎉 at NAACL 2025! Paper: aclanthology.org/2025.naacl-l... Code: github.com/lucamouchel/... #NAACL2025
A Logical Fallacy-Informed Framework for Argument Generation
Luca Mouchel, Debjit Paul, Shaobo Cui, Robert West, Antoine Bosselut, Boi Faltings. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lingu...
aclanthology.org
(NAACL) Luca and @debjit-paul.bsky.social's work looks at how we can guide LLMs to generate logically sound arguments. Introducing FIPO: a fallacy-informed framework that improves preference optimization to help LLMs avoid logical fallacies in argumentation.
Lots of great news out of the EPFL NLP lab these last few weeks. We'll be at @iclr-conf.bsky.social and @naaclmeeting.bsky.social in April / May to present some of our work in training dynamics, model representations, reasoning, and AI democratization. Come chat with us during the conference!
Translating MMLU is great, but global users of multilingual #LLMs don't care all that much about an LLM's understanding of US Law! Our new #NLProc work centers multilingual #LLM evaluations toward regional knowledge in 44 languages.
🚀 Introducing INCLUDE 🌍: A multilingual LLM evaluation benchmark spanning 44 languages! Contains *newly-collected* data, prioritizing *regional knowledge*. Setting the stage for truly global AI evaluation. Ready to see how your model measures up? #AI #Multilingual #LLM #NLProc