Joschka Strüber @Tuebingen AI Center🇩🇪

@joschkastrueber.bsky.social

PhD student at the University of Tübingen, member of @bethgelab.bsky.social, @uni_tue and @MPI_IS (IMPRS-IS). LLM multi-turn post-training and evaluations.

AI can generate correct-seeming hypotheses (and papers!). Brandolini's law states BS is harder to refute than generate. Can LMs falsify incorrect solutions? o3-mini (high) scores just 9% on our new benchmark REFUTE. Verification is not necessarily easier than generation 🧵

Bild

CuratedThoughts: Data Curation for RL Datasets 🚀 Since DeepSeek-R1 introduced reasoning-based RL, datasets like Open-R1 & OpenThoughts emerged for fine-tuning & GRPO. Our deep dive found major flaws — 25% of OpenThoughts needed elimination by data curation. Here's why 👇🧵