Nick Stracke

@rmsnorm.bsky.social

PhD Student at Ommer Lab (Stable Diffusion) Trying to understand motion... 🌐 https://nickstracke.dev

Most representation learning methods produce generic embeddings that try to preserve everything in an image. We introduce VICIS: a way to use example sets to define a tailored embedding space for what matters. This is useful whenever the desired visual signal is easier to show than to describe. 🧵👇

Source: https://poki.com/en/g/4-pics-1-word

Do we really need pixel generation to model motion? 🤔 We show how directly representing motion in a compact space enables efficient, scalable planning. 10,000× faster than video models, enabling planning and reasoning in open-world and robotics settings. Check it out ⬇️

Nick Stracke@rmsnorm.bsky.social · 4mo ago

Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇

Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇

🤔 What happens when you poke a scene — and your model has to predict how the world moves in response? We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions. It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇

Bild