CompVis - Computer Vision and Learning LMU Munich

@compvis.bsky.social

Computer Vision and Learning research group @ LMU Munich, headed by Björn Ommer. Generative Vision (Stable Diffusion, VQGAN) & Representation Learning 🌐 https://ommer-lab.com

Diffusion models treat every part of an image equally. → Same number of steps. Same compute. But images aren’t uniform. 🤔 Some regions are easy, others are hard. So why force the model to treat them the same? 🧵

Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇

🤔 What happens when you poke a scene — and your model has to predict how the world moves in response? We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions. It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇

Bild

Did you know you can distill the capabilities of a large diffusion model into a small ViT? ⚗️ We showed exactly that for a fundamental task: semantic correspondence📍 A thread 🧵👇