Thaddäus Wiedemer

@thwiedemer.bsky.social

Research Scientist at Google Deepmind Zürich | PhD student in ML at Max Planck Institute Tübingen and University of Tübingen.

How useful are self-generated 'mental images' (visual aids) in MLLM/UMM reasoning? Turns out: currently not very. Visualizations have small errors that compound in multi-step problems, and models often ignore correct visual aids in their decision making.

@jana-z.bsky.social · 6mo ago

Can AI reason by “imagining” — not just by seeing or reading? We introduce Mentis Oculi, a benchmark for machine mental imagery: multi-step visual puzzles that require maintaining and updating visual states over time. 📄 arxiv.org/abs/2602.02465 🌐 jana-z.github.io/mentis-oculi/ 🧵⬇️

🎉 Excited to present our paper VGGSounder: Audio‑Visual Evaluations for Foundation Models today at #ICCV2025! 🕦 Poster Session 1 | 11:30–13:30 📍 Poster #88 Come by if you're into audio-visual learning and want to know whether multiple modalities actually help or hurt.

Are we experiencing a 'GPT moment' in vision? In our new preprint, we show that generative video models can solve a wide range of tasks across the entire vision stack without being explicitly trained for it. 🌐 video-zero-shot.github.io 1/n

Bild

CuratedThoughts: Data Curation for RL Datasets 🚀 Since DeepSeek-R1 introduced reasoning-based RL, datasets like Open-R1 & OpenThoughts emerged for fine-tuning & GRPO. Our deep dive found major flaws — 25% of OpenThoughts needed elimination by data curation. Here's why 👇🧵