The sandwich technique came up again. So I decided to frame it properly
Scaling 4D Representations Self-supervised learning from video does scale! In our latest work, we scaled masked auto-encoding models to 22B params, boosting performance on pose estimation, tracking & more. Paper: arxiv.org/abs/2412.15212 Code & models: github.com/google-deepmind/representations4d
Generative Video Diffusion: does a model trained with this objective learn better features compared to image generation? We investigated this question and more in our latest work, please check it out! *From Image to Video: An Empirical Study of Diffusion Representations* arxiv.org/abs/2502.07001
I’m hanging out at NeurIPS this week. Come check out my co-authors’ presentations of the following Spotlight papers!