LLM leaderboards aren't passive measurements — they're mechanisms that create incentives! Why does the current evaluation protocol reward "benchmaxxing": post-training for the benchmark without improving model capability? Our #ICML2026 paper gives a theoretical answer! 🧵
Yatong Chen
@yatongchen.bsky.social
Research group leader @ Max Planck Institute working on theory & social aspect of CS. Previous @UCSC@GoogleDeepMind @Stanford @PKU1898 https://yatongchen.github.io/
At ICLR and interested in theory for LLMs? Join us at our poster to learn more about the (im)possibility of scaling laws for test-time scaling methods like Best-of-N when verification is imperfect!
Very happy to have the opportunity to share my recent work at TüWiML. Thank you for cultivating such a warm, welcoming, and nurturing community
The 4th TWiML Workshop was a blast! 🚀 With more than 80 participants, we had an engaging day with four inspiring researchers who shared their insights and experiences through presentations and a lively panel discussion. A big thank you to all participants and speakers⚡
I’ll be giving a talk at @eth-ai-center.bsky.social tomorrow, March 10, at 11:30am on LLM benchmarking incentives. Spoiler: today’s benchmarking incentives can produce unreliable model rankings, but we can fix that!
Meet me at the Benchmarking workshop (sites.google.com/view/benchma...) at EurIPS on Saturday: We’ll present two works on errors in LLM-as-Judge and their impacts on benchmarking and test-time-scaling:
Excited to be at #Neurips2025 this week to present our paper "Monoculture or Multiplicity: Which is it?", joint work with Moritz Hardt. 📄 Paper #1000: openreview.net/pdf?id=DO5Lt... 📍 Wed, Dec 3, 2025 • 4:30 PM – 7:30 PM Feel free to come by and reach out! A short 🧵.
I'll be @neuripsconf.bsky.social presenting Strategic Hypothesis Testing (spotlight!) tldr: Many high-stakes decisions (e.g., drug approval) rely on p-values, but people submitting evidence respond strategically even w/o p-hacking. Can we characterize this behavior & how policy shapes it? 1/n
We (w/ Moritz Hardt, Olawale Salaudeen and @joavanschoren.bsky.social) are organizing the Workshop on the Science of Benchmarking & Evaluating AI @euripsconf.bsky.social 2025 in Copenhagen! 📢 Call for Posters: rb.gy/kyid4f 📅 Deadline: Oct 10, 2025 (AoE) 🔗 More info: rebrand.ly/bg931sf