Super happy that our paper “Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs” just won Best Student Paper at #SIGIR2026🏆 It argues against the reflex to reach for an LLM whenever you need relevance judgments. What we found 🧵:
Lukas Gienapp
@lgnp.bsky.social
AI engineer @ Seltz; ML/IR research @ hessianAI / ScaDS.AI.
Two papers accepted at SIGIR'26 in Melbourne! Both tackle the same question: how do we scale IR evaluation reliably without compromising on human judgment as the gold standard? Spoiler: LLMs-as-a-judge does not work. 🧵1/5
We just released "German Commons", the largest openly-licensed German text dataset for LLM training: 154B tokens with clear usage rights for research and commercial use. huggingface.co/datasets/coral-nlp/german-commons
coral-nlp/german-commons · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Happy to share that our paper "The Viability of Crowdsourcing for RAG Evaluation" received the Best Paper Honourable Mention at #SIGIR2025! Very grateful to the community for recognizing our work on improving RAG evaluation. 📄 webis.de/publications...
Want to know how to make bi-encoders more than 3x faster with a new backbone encoder model? Check out our talk on the Token-Independent Text Encoder (TITE) #SIGIR2025 in the efficiency track. It pools vectors within the model to improve efficiency dl.acm.org/doi/10.1145/...
Lukas Gienapp presents "The Viability of Crowdsourcing for RAG Evaluation" at #SIGIR2025 The paper is available at: webis.de/publications...
Our paper on self-distillation for training bi-encoders got accepted at #ICTIR2025! By exploiting pretrained encoder capabilities, our approach eliminates expensive teacher models and batch sampling while maintaining the same effectiveness.
📢 Our paper "The Viability of Crowdsourcing for RAG Evaluation" has been accepted to #SIGIR2025 ! We compared how good humans and LLMs are at writing and judging RAG responses, assembling 1800+ responses across 3 styles, and 47K+ pairwise judgments in 7 quality dimensions. 🧵➡️