LLM-as-a-judge systems are often used for subjective rating tasks where humans can reasonably disagree on which rating is "correct." But how should we validate a judge system produces trustworthy ratings when humans themselves can disagree? 🧵 Paper: arxiv.org/pdf/2503.05965
Luke Guerdan
@lukeguerdan.bsky.social
PhD student @ Carnegie Mellon University I design tools and processes to support principled evaluation of AI systems. lukeguerdan.com
A subtle aspect of predictive modeling is target variable construction: the process of translating a latent, unobservable concept like "healthcare need" into a prediction target But how does target variable construction unfold in practice, and how can we better support it going forward? #CSCW2025 🧵
✨I’m on the academic job market ✨ I’m a PhD candidate at @hcii.cmu.edu studying tech, labor, and resistance 👩🏻💻💪🏽💥 I research how workers and communities contest harmful sociotechnical systems and shape alternative futures through everyday resistance and collective action More info: cella.io
Cella M. Sum –
cella.io
What can #CSCW learn from tech workers who have been involved in collective action and unionization about how to make transformative change within our field? My new #CSCW2025 paper with Mona Wang, Anna Konvicka, and Sarah Fox seeks to answer this question. Pre-print: arxiv.org/pdf/2508.12579
Have you built a generative AI evaluation that uses an LLM-as-a-judge and a rubric to rate model outputs? Sign up for a 45-minute Zoom session to provide feedback on a new tool for building trustworthy evals. Learn more at tinyurl.com/llm-as-a-judge - receive $35 for participating in a session!