When an LLM is underperforming, we tweak the prompt. Why not do the same for video models solving visual reasoning tasks? Visual prompt engineering changes the start frame appearance but not the task, improving performance. 📄 arxiv.org/pdf/2607.25537 🌐 visual-prompt-engineering.github.io
Thaddäus Wiedemer
@thwiedemer.bsky.social
Research Scientist at Google Deepmind Zürich | PhD student in ML at Max Planck Institute Tübingen and University of Tübingen.
How useful are self-generated 'mental images' (visual aids) in MLLM/UMM reasoning? Turns out: currently not very. Visualizations have small errors that compound in multi-step problems, and models often ignore correct visual aids in their decision making.
Can AI reason by “imagining” — not just by seeing or reading? We introduce Mentis Oculi, a benchmark for machine mental imagery: multi-step visual puzzles that require maintaining and updating visual states over time. 📄 arxiv.org/abs/2602.02465 🌐 jana-z.github.io/mentis-oculi/ 🧵⬇️
🚀 We're hiring! The @ellisinsttue.bsky.social leads the AI development for Germany’s new open-source nationwide Adaptive Intelligent System learning platform for schools (as part of a consortium led by Assecor & KI macht Schule, and mandated by the FWU). 👉 Apply now: forms.gle/XmLkwEDD45fY...
🎉 Excited to present our paper VGGSounder: Audio‑Visual Evaluations for Foundation Models today at #ICCV2025! 🕦 Poster Session 1 | 11:30–13:30 📍 Poster #88 Come by if you're into audio-visual learning and want to know whether multiple modalities actually help or hurt.
Are we experiencing a 'GPT moment' in vision? In our new preprint, we show that generative video models can solve a wide range of tasks across the entire vision stack without being explicitly trained for it. 🌐 video-zero-shot.github.io 1/n
Check out our newest paper! As always, it was super fun working on this with @prasannamayil.bsky.social
New preprint out! 🎉 How does LLM training loss translate to downstream performance? We show that pretraining data and tokenizer shape loss-to-loss scaling, while architecture and other factors play a surprisingly minor role! brendel-group.github.io/llm-line/ 🧵1/8
CuratedThoughts: Data Curation for RL Datasets 🚀 Since DeepSeek-R1 introduced reasoning-based RL, datasets like Open-R1 & OpenThoughts emerged for fine-tuning & GRPO. Our deep dive found major flaws — 25% of OpenThoughts needed elimination by data curation. Here's why 👇🧵
🚀 We’re hiring! Join Bernhard Schölkopf & me at @ellisinsttue.bsky.social to push the frontier of #AI in education! We’re building cutting-edge, open-source AI tutoring models for high-quality, adaptive learning for all pupils with support from the Hector Foundation. 👉 forms.gle/sxvXbJhZSccr...