Tomorrow 2pm I am giving a talk at the AfricaNLP workshop about actionable interpretability for low-resource language modelling. Stop by if you're interested in the intersection of interpretability and data-efficient modelling!
Francois Meyer
@francois-meyer.bsky.social
PhD student at the University of Cape Town, working on text generation for low-resource, morphologically complex languages. https://francois-meyer.github.io/ Cape Town, South Africa
Don’t miss the BabyBabelLM talk on Friday! Find the paper on developmentally plausible data in many languages here: aclanthology.org/2026.eacl-lo...
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Jaap Jumelet, Abdellah Fourtassi, Akari Haga, Bastian Bunzeck, Bhargav Shandilya, Diana Galvan-Sosa, Faiz Ghifari Haznitrama, Francesca Padovani, Francois Meyer, Hai Hu, Julen Etxaniz, Laurent Prevot,...
aclanthology.org
I am in Rabat for @eaclmeeting.bsky.social! I am giving an invited talk at AfricaNLP on Saturday about actionable interpretability for low-resource language modelling. @jumelet.bsky.social is presenting BabyBabelLM on Friday in the 11AM session on Parameter-Efficient Tuning and Training Dynamics.
I am in Rabat for @eaclmeeting.bsky.social! I am giving an invited talk at AfricaNLP on Saturday about actionable interpretability for low-resource language modelling. @jumelet.bsky.social is presenting BabyBabelLM on Friday in the 11AM session on Parameter-Efficient Tuning and Training Dynamics.
I am at @aaclmeeting.bsky.social in Mumbai to present our paper "The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages". I'm giving a talk tomorrow at 2pm in Room 21, in the MC-OP2 session on Multilingualism and Cross-Lingual NLP.
If a language model could dynamically optimise subword tokenisation, how would its subwords evolve during training? In our new paper we study the learning dynamics of subword segmentation: arxiv.org/pdf/2511.09197
If a language model could dynamically optimise subword tokenisation, how would its subwords evolve during training? In our new paper we study the learning dynamics of subword segmentation: arxiv.org/pdf/2511.09197
🌍Introducing BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data! LLMs learn from vastly more data than humans ever experience. BabyLM challenges this paradigm by focusing on developmentally plausible data We extend this effort to 45 new languages!
𝐃𝐨 𝐲𝐨𝐮 𝐫𝐞𝐚𝐥𝐥𝐲 𝐰𝐚𝐧𝐭 𝐭𝐨 𝐬𝐞𝐞 𝐰𝐡𝐚𝐭 𝐦𝐮𝐥𝐭𝐢𝐥𝐢𝐧𝐠𝐮𝐚𝐥 𝐞𝐟𝐟𝐨𝐫𝐭 𝐥𝐨𝐨𝐤𝐬 𝐥𝐢𝐤𝐞? 🇨🇳🇮🇩🇸🇪 Here’s the proof! 𝐁𝐚𝐛𝐲𝐁𝐚𝐛𝐞𝐥𝐋𝐌 is the first Multilingual Benchmark of Developmentally Plausible Training Data available for 45 languages to the NLP community 🎉 arxiv.org/abs/2510.10159
Today our poster will be up at @loreslm.bsky.social Poster Session #2 (2-3pm local time Abu Dhabi). It's also available online at Whova: whova.com/portal/webap...
Our paper "BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context" will be presented at The First Workshop on Language Models for Low-Resource Languages at #COLING2025 in Abu Dhabi. Paper: arxiv.org/pdf/2501.03855
Our paper "BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context" will be presented at The First Workshop on Language Models for Low-Resource Languages at #COLING2025 in Abu Dhabi. Paper: arxiv.org/pdf/2501.03855
arxiv.org