Anton Baumann

@antonbaumann.bsky.social

📢 Fully funded PhD @ Helmholtz Munich + TUM We're hiring PhD students to build trustworthy AI systems for science: explainable AI, VLMs, reliable adaptation, agentic workflows. ⏰ Priority deadline 10 Aug 2026. Application details below 👇

Bild

We introduce BayesVLM, a training-free post-hoc Bayesian method for uncertainty estimation in pretrained VLMs. BayesVLM yields interpretable, well-calibrated uncertainty with virtually no inference overhead.

Bild

SDPO enables RL agents to learn from rich feedback (i.e., not only whether an attempt failed, but why it failed, such as error messages). Even without such rich feedback, SDPO can reflect on past attempts and outperform GRPO. SDPO also accelerates solution discovery at test time!

Jonas Hübotter@jonhue.bsky.social · 6mo ago

Training LLMs with verifiable rewards uses 1bit signal per generated response. This hides why the model failed. Today, we introduce a simple algorithm that enables the model to learn from any rich feedback! And then turns it into dense supervision. (1/n)

Training LLMs with verifiable rewards uses 1bit signal per generated response. This hides why the model failed. Today, we introduce a simple algorithm that enables the model to learn from any rich feedback! And then turns it into dense supervision. (1/n)

Bild

Unfortunately, our submission to #NeurIPS didn’t go through with (5,4,4,3). But because I think it’s an excellent paper, I decided to share it anyway. We show how to efficiently apply Bayesian learning in VLMs, improve calibration, and do active learning. Cool stuff! 📝 arxiv.org/abs/2412.06014

Post-hoc Probabilistic Vision-Language Models

Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...

arxiv.org

I will present ✌️ BDU workshop papers @ NeurIPS: one by Rui Li (looking for internships) and one by Anton Baumann. 🔗 to extended versions: 1. 🙋 "How can we make predictions in BDL efficiently?" 👉 arxiv.org/abs/2411.18425 2. 🙋 "How can we do prob. active learning in VLMs" 👉 arxiv.org/abs/2412.06014

Post-hoc Probabilistic Vision-Language Models

Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...

arxiv.org

Martin Trapp@trappmartin.eurosky.social · 2y ago

On my way to #NeurIPS. Looking forward to seeing many friends again. Ping me if you want to chat, always happy to meet new people. :)