Vilém Zouhar

@zouhar.bsky.social

PhD @ ETH Zürich | working on (multilingual) evaluation of NLP | on the academic job market | go #vegan | https://vilda.net

Antropic blundered on the PR of destructivelly scanning books. Now everyone is imagining they're chopping the first prints of Mrs Dalloway but it's probably mostly "Step by Step Microsoft Access 2003". They could've gone with "We're saving books from the landfill"

Several people on socials with no technical background told me they just retweet all papers stating they don't understand what the work is about. Asked about how they could trust the paper. They said when the paper is implemented at Google, the engineers will catch the errors.

Vilém Zouhar@zouhar.bsky.social · 2mo ago

Several reviewers with no technical background told me they reviewed for NeurIPS using agents, stating they didn't understand what the work is about. Asked how they could trust the review. They said when the paper hits the socials, the people will catch the errors.

The ironic twist is that ARR *did* save us by solving our obsession with publishing papers. I know at least one person who deferred from publishing her paper because of all the annoyance&hurdles in reviewing/ACing.

me: spending 6 hours checking proofs in a paper I'm reviewing someone reviewing my paper: the paper does not evaluate whether its existence increases human evaluation adoption in practice, overall 2.5

The issue with Typst is that it's so much better than LaTeX, same as why Java was winning in corporate world. It's annoying to do things in LaTeX that go beyond the basics (text, basic styling, images, figures, tables). As a result, all papers and their tex code look the same.

Beginning to think that the reviewing disaster in comptuer science is caused by us not even liking to read papers. Some parents pay kids 1$ for each finished book. I propose we give each researcher +0.1 citations for successfully reading a paper.

Bild

Deadline extension! - The task is simple: get audio + its translation and estimate how good it is. - Mark your name as the winner of the first Speech Translation Metrics Shared Task at IWSLT 2026 🏆 Predictions submission: May 7, 2026 Description paper: May 10, 2026

Bild

I love halucinated citations in papers. They serve as an obvious canary to AI-written papers. Without them, it takes a while to notice the discourse in writing doesn't make sense or that the science is shallow or unsound.

Machine translation is tough to evaluate, partly because most of what you throw at is too easy. That doesn't at all mean that translation is solved; we're just not doing a good job finding interesting inputs.

Bild

Quality estimation (automated metrics) are amazing. Truly. We would like to use them everywhere. That gets compute-expensive very quickly. We also don't know when they don't know. In "Early-Exit and Instant Confidence Translation Quality Estimation" (at EACL26) we fix that.

Bild

Have you ever wondered how speech translation gets evaluated? Sadly, most speech evaluation downgrades to text-based metrics. Let’s do better! At IWSLT 2026, we’re launching the first-ever ✨Speech Translation Metrics Shared Task ✨!

Bild