We are super impressed by the submissions for Phase 1 of the MRL Shared Task! :) In order for us to thoroughly review the datasets, we will extend the deadline to August 14! The audio collection phase will now happen in September. Please stay tuned here (or on Discord) for more updates!
Catherine Arnett
@catherinearnett.bsky.social
NLP Researcher at EleutherAI, PhD UC San Diego Linguistics. Working on making language technologies more equitable. 📍Oxford, UK. She/her. https://catherinearnett.github.io/
Seems like a good time to share our new preprint about model openness! (with @catherinearnett.bsky.social @tylerachang.bsky.social Pamela D. Rivière, Samuel M. Taylor @camrobjones.bsky.social @seantrott.bsky.social @rplevy.bsky.social Ben Bergen, and Micah Altman): arxiv.org/abs/2603.26539
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵
This is a great opportunity for students and early-stage researchers. Please contribute if you can!
After the enthusiasm of the shared task last year, we are running another shared task to create a community-made, culturally relevant multilingual benchmark! The deadline to contribute is August 1 AoE. See more details below.
The new and expanded version of Global PIQA is out with over twice as many items. Well done to all the contributors!
We are releasing an expanded version of Global PIQA! It now covers 141 language varieties and includes parallel and non-parallel splits. We are also releasing an updated preprint.
📢 Call for Papers: 6th Multilingual Representation Learning Workshop at EMNLP in Budapest, Hungary! Join us and submit your works relating to multilingual NLP Speakers to be announced, so stay tuned! 👀 More info in the CFP: 🔗 sigtyp.github.io/ws2026-mrl.html
"[W]hen these goods remain concentrated in the hands of a few, without adequate forms of sharing and access, a new imbalance is created that contradicts the universal destination of goods" Very cool to see the Pope endorsing @eleutherai.bsky.social's mission
I’m at #HSP2026 at MIT this week! I’ll be giving a talk Friday at 5:25pm entitled “Structural Priming Effects in Language Models are Less Human-like in Languages Other Than English”. Looking forward to chatting to everyone!
@tylerachang.bsky.social and I will be presenting the Goldfish as an oral at #LREC2026 in Mallorca! 🌴
Happening now! @pjox.bsky.social and I are giving a talk for @eleutherai.bsky.social on CommonLID, a community-driven web domain evaluation dataset for language identification. Join here: discord.gg/aYy3Se7Q?eve... Paper: arxiv.org/abs/2601.18026 @commoncrawl.bsky.social
Join the EleutherAI Discord Server!
The original open science AI research collective. We started the open source LLM movement and have been pushing the boundaries of science ever since. | 33740 members
discord.gg
Announcing our latest paper: CommonLID In collaboration with @commoncrawl.bsky.social @mlcommons.org @jhu.edu we built a LID benchmark on actual Common Crawl text covering 109 languages. Existing evaluations overestimate how well LangID works on web data. arxiv.org/abs/2601.18026
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data of...
arxiv.org
Language identification still proves to be a challenging task, especially for web data. In collaboration with @mlcommons.org @eleutherai.bsky.social @jhu.edu and 97 community members, we created CommonLID, a new benchmark for LangID for 100+ languages!
We will be presenting this work this afternoon!
@jamichaelov.bsky.social and I will be presenting our paper at the CogInterp workshop 13:15 - 14:45 on Dec 7th. The paper shows how disaggregating grammatical benchmarks over the course of training reveals stages of training where models learn heuristics before learning more generalizable patterns.
I’m presenting this today at 11am. Come find me at poster #1909!
Our #NeurIPS2025 paper shows that even comparable monolingual tokenizers have different compression rates across languages. But by getting rid of whitespace tokenization and using a custom vocab size for each language, we can reduce token premiums. Preprint out now!
I’ll be in San Diego for #NeurIPS2025 next week! I will be presenting posters at the main conference and at the CogInterp workshop. I will also be at the Workshop on Evaluating AI in Practice at UCSD. Looking forward to chatting about multilingual NLP, evals, and tokenizers!
🚨 EvalEval is back - now in San Diego!🚨 🧠 Join us for the 2025 Workshop on "Evaluating AI in Practice Bridging Statistical Rigor, Sociotechnical Insights, and Ethical Boundaries" (Co-hosted with UKAISI) 📅 Dec 8, 2025 📝 Abstract due: Nov 20, 2025 Details below! ⬇️ evalevalai.com/events/works...
evalevalai.com
I’m so excited that Global PIQA is out! This has been a herculean effort by our 300+ contributors. The result is an extremely high-quality, culturally-specific benchmark for over 100 languages.
Introducing Global PIQA, a new multilingual benchmark for 100+ languages. This benchmark is the outcome of this year’s MRL shared task, in collaboration with 300+ researchers from 65 countries. This dataset evaluates physical commonsense reasoning in culturally relevant contexts.
Our #NeurIPS2025 paper shows that even comparable monolingual tokenizers have different compression rates across languages. But by getting rid of whitespace tokenization and using a custom vocab size for each language, we can reduce token premiums. Preprint out now!
In collaboration with @commoncrawl.bsky.social, MLCommons, and @eleutherai.bsky.social, the first edition of WMDQS at @colmweb.org starts tomorrow in Room 520A! We have an updated schedule on our website, including a list of all accepted papers.
I’m in Montreal this week for @colmweb.org and @wmdqs.bsky.social! Looking forward to chatting about tokenizers, multilingual data, and more! #COLM2025
I have a new blog post about the so-called “tokenizer-free” approach to language modeling and why it’s not tokenizer-free at all. I also talk about why people hate tokenizers so much!
Did you know? ❌77% of language models on @hf.co are not tagged for any language 📈For 95% of languages, most models are multilingual 🚨88% of models with tags are trained on English In a new blog post, @tylerachang.bsky.social and I dig into these trends and why they matter! 👇
We are in need of some emergency reviewers for MRL. If you are available, please fill out this form!
If you would like to sign up to be a reviewer, please fill in this form: t.co/fbunVuVhdE
We extended the deadline by one day, so you have until the end of today (Aug 24) AoE to submit! Good luck!
The deadline for MRL at #EMNLP2025 is next week! ⏰ Submission Deadline: August 23rd (AoE) 🔗 CfP: sigtyp.github.io/ws2025-mrl.h...
We have over 200 volunteers now for 90+ languages! We are hoping to expand the diversity of our language coverage and are still looking for participants who speak these languages. Check out how to get involved below, and please help us spread the word!
With six weeks left before the deadline, we have had over 50 volunteers sign up to contribute for over 30 languages. If you don’t see your language represented on the map, this is your sign to get involved!
With six weeks left before the deadline, we have had over 50 volunteers sign up to contribute for over 30 languages. If you don’t see your language represented on the map, this is your sign to get involved!
I’m in Vienna all week for @aclmeeting.bsky.social and I’ll be presenting this paper on Wednesday at 11am (Poster Session 4 in HALL X4 X5)! Reach out if you want to chat about multilingual NLP, tokenizers, and open models!
✨New pre-print✨ Crosslingual transfer allows models to leverage their representations for one language to improve performance on another language. We characterize the acquisition of shared representations in order to better understand how and when crosslingual transfer happens.
If you want to help us improve language and cultural coverage, and build an open source LangID system, please register to our shared task on Language Identification! 💬 Registering is easy! All the details are on the shared task webpage: wmdqs.org/shared-task/ Deadline: July 23, 2025 (AoE) ⏰
WMDQS: Shared Task
wmdqs.org
The Common Crawl Foundation, MLCommons, EleutherAI, and John Hopkins' Center for Language and Speech Processing have the pleasure of inviting you to register for the 1st shared task on Language Identification for web data. commoncrawl.org/blog/wmdqs-s...