Shyamgopal Karthik

@shyamgopal.bsky.social

PhD at Tübingen. Working on post-training diffusion and multimodal models. Previous research interns at Snapchat and Naver Labs. https://sgk98.github.io/

There's nothing more satisfying than watching the right noise do its magic with diffusion models! A few interesting takeaways I had from this work 🧵

A. Sophia Koepke@askoepke.bsky.social · 2mo ago

#CVPR2026 paper: It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models Text-to-image models often collapse to near-identical samples. Our fix: optimize the noise. Start from pink 🩷, not white noise. 🔗 akoepke.github.io/divgen/index... 1/6

New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

akoepke.github.io

🚨 New Paper: "Solving Spatial Supersensing Without Spatial Supersensing" Huge credit to the Cambrian-S team for tackling one of the hardest open problems in video understanding: spatial supersensing. In our paper, we take a closer look at their benchmarks & methods 👇

Bild

Unfortunately, our submission to #NeurIPS didn’t go through with (5,4,4,3). But because I think it’s an excellent paper, I decided to share it anyway. We show how to efficiently apply Bayesian learning in VLMs, improve calibration, and do active learning. Cool stuff! 📝 arxiv.org/abs/2412.06014

Post-hoc Probabilistic Vision-Language Models

Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...

arxiv.org

I've been talking about writing this paper to anyone who would listen since 2020. I bombed a bunch of job talks trying to convince companies to work on this. It's so nice to finally just be able to say, yes, self-play RL in a diverse world gives you immense capabilities arxiv.org/abs/2502.03349

Robust Autonomy Emerges from Self-Play

Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic drivi...

arxiv.org

This is maybe my favorite thing I've seen out of #NeurIPS2024. Head over to HuggingFace and play with this thing. It's quite extraordinary.

Luca Eyring@lucaeyring.bsky.social · 2y ago

Thanks to @fffiloni.bsky.social and @natanielruiz.bsky.social, we have a running live Demo of ReNO, play around with it here: 🤗: huggingface.co/spaces/fffil... We are excited to present ReNO at #NeurIPS2024 this week! Join us tomorrow from 11am-2pm at East Exhibit Hall A-C #1504!

Can we enhance the performance of T2I models without any fine-tuning? We show that with our ReNO, Reward-based Noise Optimization, one-step models consistently surpass the performance of all current open-source Text-to-Image models within the computational budget of 20-50 sec! #NeurIPS2024

Bild

I will present ✌️ BDU workshop papers @ NeurIPS: one by Rui Li (looking for internships) and one by Anton Baumann. 🔗 to extended versions: 1. 🙋 "How can we make predictions in BDL efficiently?" 👉 arxiv.org/abs/2411.18425 2. 🙋 "How can we do prob. active learning in VLMs" 👉 arxiv.org/abs/2412.06014

Post-hoc Probabilistic Vision-Language Models

Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...

arxiv.org

Martin Trapp@trappmartin.eurosky.social · 2y ago

On my way to #NeurIPS. Looking forward to seeing many friends again. Ping me if you want to chat, always happy to meet new people. :)

After a break of over 2 years, I'm attending a conference again! Excited to attend NeurIPS, even more so to be presenting ReNO, getting inference-time scaling and preference optimization to work for text-to-image generation. Do reach out if you'd like to chat!

Bild

🚨New Paper Alert🚨 🚀 Introducing FlowChef, "Steering Rectified Flow Models in the Vector Field for Controlled Image Generation"! 🌌✨ - Perform image editing, solve inverse problems, and more. - Achieved inversion-free, gradient-free, & training-free inference time steering! 🤯 👇👇

Bild