Jindřich Libovický

@jlibovicky.bsky.social

Researcher at Charles University | multilingual natural language processing, machine translation

Very proud of my student Lukáš Eigler presenting this at the ACL Student Research Workshop. TL;DR: you can validate NLP evaluation metrics with synthetic LLM judgments — the rankings track human ones almost perfectly. 📄 aclanthology.org/2026.acl-srw.125

LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation

Lukáš Eigler, Jindřich Libovický, David Hurych. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). 2026.

aclanthology.org

Institute of Formal and Applied Linguistics@ufal.mff.cuni.cz · 4w ago

LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation aclanthology.org/2026.acl-srw... by Lukáš Eigler, @jlibovicky.bsky.social & David Hurych Rankings from synthetic LLM-generated data almost perfectly match real human judgments. 🤖⚖️ at Student Research Workshop

Dušan made something like TensorBoard for tokenization 📊 Token length, entropy, vocab overlap, JS divergence, alignment scores: sliced by language family, script, region, speakers, data availability. 🔗 Code: github.com/ufal/TokCollate 🔗 Demo: quest.ms.mff.cuni.cz/tokcollate

BildBild
Institute of Formal and Applied Linguistics@ufal.mff.cuni.cz · last mo.

Dušan Variš presents joint work w/ @abyste.bsky.social & @jlibovicky.bsky.social: TokCollate: A Comprehensive Tool for Tokenizer Evaluation and Visualization across Languages aclanthology.org/2026.acl-dem... A dashboard for your tokenizer 👀 See which languages you left behind. Open-source, MIT.

AC observation this ARR cycle: I'm getting way more reviews and way earlier, than usual. I suspect it's the draconian desk-reject threats working, and not a sign of better reviewing culture. I worry that this gets read as 'threats work, do more of them, discipline the lazy researchers'.

Nice work by my student @gianlucavico.bsky.social on a topic close to home: crowdsourcing Piedmontese to test LLMs on non-standard orthography. New dataset covering tokenization, classification & translation.

Institute of Formal and Applied Linguistics@ufal.mff.cuni.cz · 4mo ago

Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography by @gianlucavico.bsky.social and @jlibovicky.bsky.social aclanthology.org/2026.vardial... New Piedmontese dataset tests tokenization, classification & translation! 🗣️

So proud of my new PhD student Adnan Al Ali for presenting their master's thesis work at EACL! 🎓 A great contribution to understanding bias in AI text detectors across languages.

Institute of Formal and Applied Linguistics@ufal.mff.cuni.cz · 4mo ago

Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors by Adnan Al Ali, @jindrahelcl.bsky.social and @jlibovicky.bsky.social aclanthology.org/2026.eacl-sr... AI text detectors are suspected of biasing against non-native speakers, but not in Czech 🇨🇿

I'm at #EACL2026 in Rabat 🇲🇦. Find me and talk to me about tokenization or multilingual model eval. Also, check out our work on how to eval morphological plausibility of your tokenizer if you don't have gold segmentation data, but you happen to have morphosyntactic features 👇

Institute of Formal and Applied Linguistics@ufal.mff.cuni.cz · 4mo ago

Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features by @abyste.bsky.social & @jlibovicky.bsky.social aclanthology.org/2026.finding... TL;DR: Morpho-syntax features replace gold segmentation data for tokenization eval.

Spent time making AI-generated images of Bayes' Rule, Laplace Smoothing, Markov Chains & Shannon Entropy for class today 🎨🤖 Even though the images are objectively hilarious, none of the 50 students in the room laughed. Or even smiled. 💀

Bild

So proud of my PhD student @andrei-a-manea.bsky.social for his first first-author publication! 🎉 He presented this work last week at TSD. Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders arxiv.org/pdf/2504.21681

arxiv.org

Institute of Formal and Applied Linguistics@ufal.mff.cuni.cz · 11mo ago

Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders by @andrei-a-manea.bsky.social & @jlibovicky.bsky.social TL;DR: Explore how parallel datasets improve cross-lingual transfer in vision-language models. arxiv.org/abs/2504.21681

🧵 We're releasing CUS-QA - a new benchmark for testing LLMs on regional knowledge! Find out what your model knows about Czechia 🇨🇿, Slovakia 🇸🇰, and Ukraine 🇺🇦! 👉 Textual and visual questions, answers, and human judgment on model outputs! huggingface.co/datasets/ufa... www.arxiv.org/abs/2507.22752

ufal/cus-qa · Datasets at Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

Just presented MAGBIG, a new dataset and evaluation methodology for gender bias in multilingual text-to-image generation. Grammatical gender matters when studying these biases across languages! Thanks to Felix Friedrich, @kathaem.bsky.social and all co-authors - it was fun to work on this together!

Institute of Formal and Applied Linguistics@ufal.mff.cuni.cz · last yr.

Multilingual Text-to-Image Generation Magnifies Gender Stereotypes aclanthology.org/2025.acl-lon... by Felix Friedrich, @kathaem.bsky.social, Patrick Schramowski, @mbrackaiml.bsky.social , @jlibovicky.bsky.social, @kerstingaiml.bsky.social, Alex Fraser