July has been a good month: * Our ACL paper about Wikipedia quality was awarded an SAC highlight (aclanthology.org/2026.acl-lon...) * @colemanhaley.bsky.social joined our lab as a postdoc * Coleman's CoNLL paper about impossible languages won the best paper award (aclanthology.org/2026.conll-m...)
LAGoM NLP
@lagom-nlp.bsky.social
We are the Leuven AI Group of Multilingual NLP (LAGoM NLP), a research lab at the department of Computer Science at KU Leuven, led by @mdlhx
This Sunday at 5pm at #ACL2026 in the multilingual session, Kushal Tatariya and Artur Kulmizev will present our work on auditing the quality of Wikipedia for low-resource NLP, see the paper here: aclanthology.org/2026.acl-lon...
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, Miryam de Lhoneux. Proceedings of the 64th Annual Meeting of the Associati...
aclanthology.org
BPE-knockout just got outperformed by an algorithm that modifies BPE tokenisers in a feedback loop to make them absorb more and more constraints. It doesn't even need more data to do that. It uses the tokeniser itself as a dataset. 🧵
LAGoM will present several papers at #EACL2026 in Rabat next week! Our work at this year’s conference spans tokenisation, multilingual evaluation, and model design.
New EACL paper (with @mdlhx.bsky.social)! We tested if comparing perplexity of parallel data across languages is fair. Turns out: it depends. We show the choice of test set (even with consistent meaning) can flip conclusions about which language is easier to model. Paper: arxiv.org/abs/2601.10580
Form and Meaning in Intrinsic Multilingual Evaluations
Intrinsic evaluation metrics for conditional language models, such as perplexity or bits-per-character, are widely used in both mono- and multilingual settings. These metrics are rather straightforwar...
arxiv.org
When is a language hard to model? Previous research has suggested that morphological complexity both does and does not play a role, but it does so by relating the performance of language models to corpus statistics of words or subword tokens in isolation.
Our group has two papers at #acl2025: * (Main) GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model by Thomas Bauwens, David Kaczér and @mdlhx.bsky.social, presented by Thomas. URL: aclanthology.org/2025.acl-lon...
GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model
Thomas Bauwens, David Kaczér, Miryam De Lhoneux. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
aclanthology.org
The submission deadline for #CLIN35 has been extended by one week! New deadline: June 20th. 🔊 Spread the word! More info: clin35.ccl.kuleuven.be/call-for-abs...
CLIN35 - Call for Abstracts
We invite submissions for CLIN35, the 35th edition of the Computational Linguistics in the Netherlands (CLIN) conference, which will take place in Leuven on September 12th, 2025. Abstracts describing ...
clin35.ccl.kuleuven.be
Reminder, a few more days to apply!
Interested in multilingual tokenization in #NLP? Lisa Beinborn and I are hiring! PhD candidate position in Göttingen, Germany: www.uni-goettingen.de/de/644546.ht... PostDoc position in Leuven, Belgium: www.kuleuven.be/personeel/jo... Deadline 6th of June
📅 Don't forget! The deadline for submitting your abstract to the #CLIN conference in Leuven is coming: 13th of June! Submitting is easy: name, title of your work, 500-word abstract, done! #nlp #nlproc #compling #llm #ai #dutch clin35.ccl.kuleuven.be
CLIN35
Computational Linguistics in The Netherlands (CLIN) is a yearly conference on computational linguistics. Each year the conference is organized by a different institution in the Dutch-speaking region. ...
clin35.ccl.kuleuven.be
We are hiring in #nlproc!!
Interested in multilingual tokenization in #NLP? Lisa Beinborn and I are hiring! PhD candidate position in Göttingen, Germany: www.uni-goettingen.de/de/644546.ht... PostDoc position in Leuven, Belgium: www.kuleuven.be/personeel/jo... Deadline 6th of June
I’m looking for a postdoc, to start ideally ASAP! The work would be in the EU-funded TrustLLM project, focusing on modularisation and language adaptation of LLMs, tokenization, and evaluation benchmarks for multilingual LLMs. The position would be full-time for 2 years with no teaching obligation.
@wpoelman.bsky.social and @mdlhx.bsky.social 's 🔥 hot takes on multilingual LLM evaluation, to appear @nodalida.bsky.social is up on arXiv: arxiv.org/abs/2412.08392
The Roles of English in Evaluating Multilingual Language Models
Multilingual natural language processing is getting increased attention, with numerous models, benchmarks, and methods being released for many languages. English is often used in multilingual evaluati...
arxiv.org
🚨 New Account Alert! This is the official account of the *MilaNLP group*. We had to recreate it because it was not indexed. If you were following us before, please follow us again. If not, now’s the perfect time to start!
There's too many starter packs. 👇 Here's a list, mostly for NLP, ML, and related areas.
NLP grad students
Join the conversation
go.bsky.app
Our work on quality estimation of non-English Wikipedia articles is on arXiv! 🎉 arxiv.org/abs/2411.055... By Kushal Tatariya, Artur Kulmizev, Wessel Poelman, @estherploeger.bsky.social, @bollmann.me, Johannes Bjerva, Jiaming Luo, Heather Lent and @mdlhx.bsky.social ✨
How Good is Your Wikipedia?
Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in multilingual NLP. In the context of low-resource languages, however, these quality assum...
arxiv.org
The NLP labs starter pack is here! go.bsky.app/LKGekew Let us know if you want to be added!
Our team at #EMNLP2024 this week!
Our EMNLP papers are out on the arXiv! 2 main conference papers, about pixology: arxiv.org/abs/2410.12011, about typological diversity: arxiv.org/abs/2402.04222 and one workshop paper about zero-shot pos tagging: arxiv.org/abs/2410.10576. Looking forward to connecting in Miami!
Our EMNLP papers are out on the arXiv! 2 main conference papers, about pixology: arxiv.org/abs/2410.12011, about typological diversity: arxiv.org/abs/2402.04222 and one workshop paper about zero-shot pos tagging: arxiv.org/abs/2410.10576. Looking forward to connecting in Miami!
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models
Pixel-based language models have emerged as a compelling alternative to subword-based language modelling, particularly because they can represent virtually any script. PIXEL, a canonical example of su...
arxiv.org
New #NLProc paper on ArXiv by E. Ploeger, W. Poelman, M. de Lhoneux and J. Bjerva! More and more papers in NLP claim to evaluate on ‘typologically diverse’ languages. But what does this even mean? We systematically investigate such claims. arxiv.org/abs/2402.04222
New paper: 'Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification' t.co/MpOmz4uUsd Kushal will present it at SIGTYP @ EACL 2024! #NLProc
Hello bluesky! We have a website and a bsky account, therefore we exist! (We are also in the bad place but will try to always post on here first)