Come meet me at the last poster session at ICML if you are interested in chatting about LLM evaluation and pairwise comparisons! I will be presenting joint work with my advisor (Moritz Hardt). Thu Jul 9, 5:00 PM – 6:45 PM KST, Hall A (#4411) arxiv.org/abs/2606.09409
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge bi...
arxiv.org