Mila Gorecki

@milago.bsky.social

PhD student in Machine Learning @ MPI-IS Tübingen, Tübingen AI Center, IMPRS-IS

Takeaway: leaderboard design is mechanism design! If you care about evals, post-training, or strategic behavior in ML, come talk to us! 📍Poster: HALL A #4413 📅 Thursday, July 9th, 2:30pm 📄 Paper link: arxiv.org/abs/2603.08371 Joint work w/ Guanhua Zhang & Moritz Hardt

Leaderboard Incentives: Model Rankings under Strategic Post-Training

Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmaxxing or training on ...

arxiv.org

LLM leaderboards aren't passive measurements — they're mechanisms that create incentives! Why does the current evaluation protocol reward "benchmaxxing": post-training for the benchmark without improving model capability? Our #ICML2026 paper gives a theoretical answer! 🧵

Bild

Come meet me at the last poster session at ICML if you are interested in chatting about LLM evaluation and pairwise comparisons! I will be presenting joint work with my advisor (Moritz Hardt). Thu Jul 9, 5:00 PM – 6:45 PM KST, Hall A (#4411) arxiv.org/abs/2606.09409

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge bi...

arxiv.org

At #NeurIPS in San Diego this week? Interested in XAI, causality, or performative prediction? Come visit our poster! 💬 Performative Validity of Recourse Explanations 📆 Wednesday, 4.30 pm, Poster Session 2 w/ Hidde Fokkema, Timo Freiesleben, Celestine Mendler-Dünner, Ulrike von Luxburg

Bild