Lukas Gienapp

@lgnp.bsky.social

AI engineer @ Seltz; ML/IR research @ hessianAI / ScaDS.AI.

Super happy that our paper “Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs” just won Best Student Paper at #SIGIR2026🏆 It argues against the reflex to reach for an LLM whenever you need relevance judgments. What we found 🧵:

BildBild

Two papers accepted at SIGIR'26 in Melbourne! Both tackle the same question: how do we scale IR evaluation reliably without compromising on human judgment as the gold standard? Spoiler: LLMs-as-a-judge does not work. 🧵1/5

Bild

Our paper on self-distillation for training bi-encoders got accepted at #ICTIR2025! By exploiting pretrained encoder capabilities, our approach eliminates expensive teacher models and batch sampling while maintaining the same effectiveness.

Bild

📢 Our paper "The Viability of Crowdsourcing for RAG Evaluation" has been accepted to #SIGIR2025 ! We compared how good humans and LLMs are at writing and judging RAG responses, assembling 1800+ responses across 3 styles, and 47K+ pairwise judgments in 7 quality dimensions. 🧵➡️

Bild