Vilém Zouhar
@zouhar.bsky.social
PhD @ ETH Zürich | working on (multilingual) evaluation of NLP | on the academic job market | go #vegan | https://vilda.net
Eine deutsche KI behauptet, Schweizerdeutsch zu können. Unser neuer Test sagt: «Chabis». vamvas.ch/benchmark-sw...
Eine deutsche KI behauptet, Schweizerdeutsch zu können. Unser neuer Test sagt: «Chabis»
vamvas.ch
Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for ...
arxiv.org
Antropic blundered on the PR of destructivelly scanning books. Now everyone is imagining they're chopping the first prints of Mrs Dalloway but it's probably mostly "Step by Step Microsoft Access 2003". They could've gone with "We're saving books from the landfill"
Last chance (3 days!) to participate in the refreshed WMT evaluation shared task. Pick any subtask and submit an automatic metric that's aligned with human annotations of translation quality across many languages. www2.statmt.org/wmt26/mteval...
Several people on socials with no technical background told me they just retweet all papers stating they don't understand what the work is about. Asked about how they could trust the paper. They said when the paper is implemented at Google, the engineers will catch the errors.
Several reviewers with no technical background told me they reviewed for NeurIPS using agents, stating they didn't understand what the work is about. Asked how they could trust the review. They said when the paper hits the socials, the people will catch the errors.
Several reviewers with no technical background told me they reviewed for NeurIPS using agents, stating they didn't understand what the work is about. Asked how they could trust the review. They said when the paper hits the socials, the people will catch the errors.
😭
Not many people know this but this button was not invented to complain that reviewers didn't give you a higher score.
There are many things machine translation still can't do. Help us steer the next direction by contributing hard-to-translate inputs (and be on a cool paper).
The best rebuttal I read is just 4 sentences in 3 bullet points clarifying my listed weaknesses.
All my homies left for ACL/ICML and left me home managing a project *for which we're looking for coauthors* 🥸. Join us (also talk to @sethjsa.bsky.social @onadegibert.bsky.social @niyatibafna.bsky.social @patuchen.bsky.social @michellewastl.bsky.social @ayukh.bsky.social)
The ironic twist is that ARR *did* save us by solving our obsession with publishing papers. I know at least one person who deferred from publishing her paper because of all the annoyance&hurdles in reviewing/ACing.
Be the reviewer you want (or your AC wants) to have. (The 1 goes to the AI written paper. No thank you for wasting 2 hours of my life. Pleasure reviewing the rest.)
me: spending 6 hours checking proofs in a paper I'm reviewing someone reviewing my paper: the paper does not evaluate whether its existence increases human evaluation adoption in practice, overall 2.5
The issue with Typst is that it's so much better than LaTeX, same as why Java was winning in corporate world. It's annoying to do things in LaTeX that go beyond the basics (text, basic styling, images, figures, tables). As a result, all papers and their tex code look the same.
Beginning to think that the reviewing disaster in comptuer science is caused by us not even liking to read papers. Some parents pay kids 1$ for each finished book. I propose we give each researcher +0.1 citations for successfully reading a paper.
We are at @eamt2026.bsky.social with a large-scale stealth project called "[redacted] Translation [redacted]". Looking for contributors across all languages to challenge MTs. Talk to @patuchen.bsky.social, @sethjsa.bsky.social and me to be a coauthor!
We'll be hosting a tutorial at EAMT (already next week!), KONVENS and MT Marathon on human evaluation. Come learn with us! With support of @maikezufle.bsky.social and @patuchen.bsky.social
I reviewed for ICML and all I got was this lousy registration.
Deadline extension! - The task is simple: get audio + its translation and estimate how good it is. - Mark your name as the winner of the first Speech Translation Metrics Shared Task at IWSLT 2026 🏆 Predictions submission: May 7, 2026 Description paper: May 10, 2026
I love halucinated citations in papers. They serve as an obvious canary to AI-written papers. Without them, it takes a while to notice the discourse in writing doesn't make sense or that the science is shallow or unsound.
Come to La Palmaraie (EACL) for the first Multilingual Multicultural Evaluation workshop! 🧐 now. Organized by @pinzhen.bsky.social @hanxuhu.bsky.social @simi97k.bsky.social Wenhao Zhu @bazril.bsky.social Alexandra Birch @afaji.bsky.social Rico Sennrich @sarahooker.bsky.social
saddest conversation at eacl: - what do you work on? - mathematical modelling of evaluation - oh what kind of LLM is that?
all conference attendees under the age of 12 agree that playing subway surfer vastly improves the poster presentation experience
How Important is ‘Perfect’ English for Machine Translation Prompts? by @patuchen.bsky.social, @niyatibafna.bsky.social, @sethjsa.bsky.social , @gianlucavico.bsky.social ,W. Kamzela, @kathaem.bsky.social & @zouharvi.bsky.social aclanthology.org/2026.finding... TL;DR: Prompt errors < prompt choice
Machine translation is tough to evaluate, partly because most of what you throw at is too easy. That doesn't at all mean that translation is solved; we're just not doing a good job finding interesting inputs.
Quality estimation (automated metrics) are amazing. Truly. We would like to use them everywhere. That gets compute-expensive very quickly. We also don't know when they don't know. In "Early-Exit and Instant Confidence Translation Quality Estimation" (at EACL26) we fix that.
How often is human evaluation skipped in papers/workflows just because it's too difficult to set up? Yet even small humeval can give so much more signal than automatic metrics. Introducing Pearmut, Human Evaluation of Translation Made Trivial🍐 arxiv.org/pdf/2601.02933
Have you ever wondered how speech translation gets evaluated? Sadly, most speech evaluation downgrades to text-based metrics. Let’s do better! At IWSLT 2026, we’re launching the first-ever ✨Speech Translation Metrics Shared Task ✨!