🧐 Can you distill knowledge from 4 vision teachers into one student using ZERO real images? 🥳 Turns out: YES, and surprisingly well. 📣 Our ECCV 2026 paper "IDeaL" closes most of the gap with real-image distillation, using optimized structured noise.
Mert Bulent Sariyildiz
@mbsariyildiz.bsky.social
Research scientist at Naver Labs Europe. https://mbsariyildiz.github.io/
1/6 Excited to share that our paper on model merging was accepted at ECCV 2026! 🎉 We introduce an efficient, decoder-free proxy that makes model selection faster, simpler and practical across vision tasks. 📄 arxiv.org/abs/2604.12935 🌐 europe.naverlabs.com/task-alignment 🧵👇
For your Embodied AI task you want a recurrent model with constant complexity per step, but you don't want to lose the power of transformers (which store the full obs history and attend to it)? Do not despair, we have your back. We distill transformers into recurrent transformers 1/8
Wanna the outstanding performance of MASt3R while using a ViT-B or ViT-S encoder instead of its ViT-L one? Don't miss how we build DUNE, a single encoder for diverse 2D & 3D tasks, at this afternoon #CVPR2025 poster session (poster #376). paper: arxiv.org/abs/2503.14405 code: github.com/naver/dune
1/ 📄 Paper 1: "DUNE: Distilling a UNiversal Encoder from Heterogeneous 2D and 3D Teachers" We propose DUNE: a ViT-based encoder distilled from multiple specialized 2D & 3D foundation models to unify visual tasks across 2D, 3D and human understanding.
🧵 Two new papers at #CVPR2025 on generalization in visual representations — covering both universal encoders for 2D and 3D tasks, and open-vocabulary semantic segmentation Let's dive in! 👇