Takeaway: leaderboard design is mechanism design! If you care about evals, post-training, or strategic behavior in ML, come talk to us! 📍Poster: HALL A #4413 📅 Thursday, July 9th, 2:30pm 📄 Paper link: arxiv.org/abs/2603.08371 Joint work w/ Guanhua Zhang & Moritz Hardt
Leaderboard Incentives: Model Rankings under Strategic Post-Training
Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmaxxing or training on ...
arxiv.org