I've already collected my first two stickers for the @ecir2026.eu Collab-a-thon. If you're in Delft today, don't muss out the first collab-a-thon session at 4pm in LAB.115 🤝 #collab-a-thon #collaboration #research #ecir
Jan Heinrich Merker
@heinrich.merker.id
📚 Researcher • 💻 Developer • 🇪🇺 European PhD student for health-related information retrieval at @uni-jena.de × @webis.de
We just released "German Commons", the largest openly-licensed German text dataset for LLM training: 154B tokens with clear usage rights for research and commercial use. huggingface.co/datasets/coral-nlp/german-commons
coral-nlp/german-commons · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Honored to win the ICTIR Best Paper Honorable Mention Award for "Axioms for Retrieval-Augmented Generation"! Our new axioms are integrated with ir_axioms: github.com/webis-de/ir_... Nice to see axiomatic IR gaining momentum.
We presented two papers at ICTIR 2025 today: - Axioms for Retrieval-Augmented Generation webis.de/publications... - Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins webis.de/publications...
Happy to share that our paper "The Viability of Crowdsourcing for RAG Evaluation" received the Best Paper Honourable Mention at #SIGIR2025! Very grateful to the community for recognizing our work on improving RAG evaluation. 📄 webis.de/publications...
Lets replace search with "AI" then! Totally logical if you ask me. Even more worth it when you know they're exponentially overtaking the airline industry in their carbon footprint. Study: www.cjr.org/tow_center/w...
Decades from now, the Covid-19 pandemic will be visible in the historical data of nearly anything measurable today. Here’s an incomplete collection of charts that capture that break — across the economy, health care, education, work, family life and more.
30 Charts That Show How Everything Changed in March 2020
It can be easy to forget, or look away from, the pain and disruption of the pandemic. The numbers will be there to remind us.
nytimes.com
New preprint of WSDM demo by @maik_froebe @matthias and Ferdinand Schlatt Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval https://arxiv.org/abs/2411.04677 https://webis.de/lightning-ir/
Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
A wide range of transformer-based language models have been proposed for information retrieval tasks. However, including transformer-based models in retrieval pipelines is often complex and requires substantial engineering effort. In this paper, we introduce Lightning IR, an easy-to-use PyTorch Lightning-based framework for applying transformer-based language models in retrieval scenarios. Lightning IR provides a modular and extensible architecture that supports all stages of a retrieval pipeline: from fine-tuning and indexing to searching and re-ranking. Designed to be scalable and reproducible, Lightning IR is available as open-source: https://github.com/webis-de/lightning-ir.
arxiv.org
What a team of keynote speakers. I must confess seeing that Steve Robertson will be there is a thrill. One of the legends of information retrieval reflecting on the field. #sigir2025 sigir2025.dei.unipd.it/keynote-spea...
SIGIR 2025, Padua, 13-18 July | Keynotes
The SIGIR 2025 keynotes are held by esteemed speakers: Robertson S., Gurevych I. and Frieder O., who will cover topics that range from AI in medical search and ecommendation to BM25 and probabilistic ...
sigir2025.dei.unipd.it
🚨 New Pre-Print! 🚨 Reviewer 2 has once again asked for DL’19, what can you say in rebuttal? To help, we have re-annotated DL’19. Work done with @maik_froebe.bsky.social, @hscells.bsky.social, @fschlatt1.bsky.social, Guglielmo Faggioli, Saber Zerhoudi, @macavaney.bsky.social, Eugene Yang 🧵
Andrew Parry, Maik Fr\"obe, Harrisen Scells, Ferdinand Schlatt, Guglielmo Faggioli, Saber Zerhoudi, Sean MacAvaney, Eugene Yang Variations in Relevance Judgments and the Shelf Life of Test Collections https://arxiv.org/abs/2502.20937
I'm putting together a slide illustrating how generative AI is being forced on people even though they don't want it, and this is sort of funny. Here's are Google's autocomplete suggestions for "google gemini how to", and Bing's autocomplete suggestions for for "microsoft copilot how to".
Excited to have received 3k stars on GitHub! 🎉 github.com/janheinrichm... Some stats: ⭐ 3,000 stars 🔀 579 forks 👁️ 413 followers 📚 108 public repositories 📖 101 open-source licensed 💾 3.7 GB of source code ✏️ 10,217 commits 💬 373 issues 🚀 233 pull requests Thanks to all stargazers and followers! ☺️
Great first day of #TREC2024! Especially the panel on evaluations of RAG approaches was very insightful 👍 Excited to see GenIR evaluation getting more and more solid 🙂
Actually, we already submitted some very similar approaches to TREC BioGen. Let's see how that plays out 😉
Follow-up on our #BIOASQ2024 submission: We actually submitted the best approach for some of the tasks 👍 Looking forward to further improving Medical RAG! #CLEF2024
https://www.theguardian.com/science/2024/feb/03/the-situation-has-become-appalling-fake-scientific-papers-push-research-credibility-to-crisis-point?CMP=Share_iOSApp_Other
Reminder that research that relies on OpenAI models is (usually) not reproducible.
📄 Pre-print: https://webis.de/publications.html?q=argument#reimer_2023b 💾 Code: https://t.co/FKq6Yx4cto
Hence, we propose better few-shot and zero-shot stance detectors based on GPT-3.5 and Flan-T5. Our GPT-3.5 stance detector reaches an F1 of 0.49 and is able to push the top-3 systems of Touché to the top of the leaderboard ⏫
Our short paper “Stance-Aware Re-Ranking for Non-factual Comparative Queries” with @albondarenko2, @maik_froebe, and @matthias_hagen got accepted at #ArgMininig 2023 ☺️ Takeaway: Improve nDCG by moving docs that take no stance down the result list. @ArgminingOrg #EMNLP #NLProc
With Vienna conveniently located at the center of the European rail network, I'm taking the night train via Berlin. Safe and sustainable travels, everyone! ☺️
Now I'm also headed to @essir_eu 🚆 I'm looking forward to a week full of hands-on IR courses – especially on Saturday, where there will be several lectures about health-related IR! #essir2023
What a week 😄 1️⃣ Our resource paper is accepted at @SIGIRConf #sigir2023 📄🔍 2️⃣ I submitted my Master's thesis 🎉☺️ Now it's time for vacation until I join @webis_de as a PhD student in May 👍
The other direction is more problematic: how can we as researchers use models/code/ideas from companies that don't publish the underlying concepts? GPT-4 etc. are effectively black boxes and should therefore be used very carefully.
I don't agree. We're doing science for the public and that includes companies as well. If you don't want your models/code/ideas to be used by anyone, then write a patent instead of a paper. Or release your model weights and data under some less permissive license.
I'm very honored to be awarded the #Deutschlandstipendium from @BMBF_Bund and @LBBW at @UniHalle. Thanks for supporting my continuing research on web search and information retrieval through this academic scholarship!
Little #LifeHack for everyone who uses @SumUp Accounting and is currently filing the German annual VAT report (USt-Jahresanmeldung) 💵🇩🇪 This link gives you an annual report of all VAT-related transactions: https://accounting.sumup.com/reports/vat/period?fromDate=2021-01-01&toDate=2021-12-31 For othe
Do you have any other favorite IntelliJ plugins, especially for computer scientists and data scientists?
Here are 5 outstanding IntelliJ plugins I use every day and think are worth checking out for data scientists. 🧵 #IntelliJ @intellijidea