Very proud of my student Lukáš Eigler presenting this at the ACL Student Research Workshop. TL;DR: you can validate NLP evaluation metrics with synthetic LLM judgments — the rankings track human ones almost perfectly. 📄 aclanthology.org/2026.acl-srw.125
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
Lukáš Eigler, Jindřich Libovický, David Hurych. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). 2026.
aclanthology.org
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation aclanthology.org/2026.acl-srw... by Lukáš Eigler, @jlibovicky.bsky.social & David Hurych Rankings from synthetic LLM-generated data almost perfectly match real human judgments. 🤖⚖️ at Student Research Workshop