Thomas Wimmer

@wimmerthomas.bsky.social

PhD Candidate at the Max Planck ETH Center for Learning Systems working on 3D Computer Vision. https://wimmerth.github.io

DUSt3R et al. are impressive, but how do they actually work? We investigate this in our project ๐˜œ๐˜ฏ๐˜ฅ๐˜ฆ๐˜ณ๐˜ด๐˜ต๐˜ข๐˜ฏ๐˜ฅ๐˜ช๐˜ฏ๐˜จ ๐˜”๐˜ถ๐˜ญ๐˜ต๐˜ช-๐˜๐˜ช๐˜ฆ๐˜ธ ๐˜›๐˜ณ๐˜ข๐˜ฏ๐˜ด๐˜ง๐˜ฐ๐˜ณ๐˜ฎ๐˜ฆ๐˜ณ๐˜ด!โฃ We share findings on the iterative nature of reconstruction, the roles of cross and self-attention, and the emergence of correspondences across the network [1/8] โฌ‡๏ธ

Vision and Graphics Trends@si-cv-graphics.bsky.social ยท 9mo ago

๐—จ๐—ป๐—ฑ๐—ฒ๐—ฟ๐˜€๐˜๐—ฎ๐—ป๐—ฑ๐—ถ๐—ป๐—ด ๐— ๐˜‚๐—น๐˜๐—ถ-๐—ฉ๐—ถ๐—ฒ๐˜„ ๐—ง๐—ฟ๐—ฎ๐—ป๐˜€๐—ณ๐—ผ๐—ฟ๐—บ๐—ฒ๐—ฟ๐˜€ Michal Stary, Julien Gaubil, Ayush Tewari, Vincent Sitzmann arxiv.org/abs/2510.24907 Trending on www.scholar-inbox.com

๐Ÿค” What if you could generate an entire image using just one continuous token? ๐Ÿ’ก It works if we leverage a self-supervised representation! Meet RepTok๐ŸฆŽ: A generative model that encodes an image into a single continuous latent while keeping realism and semantics. ๐Ÿงต ๐Ÿ‘‡

Bild

Suppose you have separate datasets X, Y, Z, without known correspondences. We do the simplest thing: just train a model (e.g., a next-token predictor) on all elements of the concatenated dataset [X,Y,Z]. You end up with a better model of dataset X than if you had trained on X alone! 6/9

Architecture for Unpaired Multimodal Learner.

Happy to find my name on the list of outstanding reviewers :] Come and check out our poster on learning better features for semantic correspondence in Hawaii! ๐Ÿ“ Poster #538 (Session 2) ๐Ÿ—“๏ธ Oct 21 | 3:15 โ€“ 5:00 p.m. HST genintel.github.io/DIY-SC

#ICCV2025@iccv.bsky.social ยท 10mo ago

Thereโ€™s no conference without the efforts of our reviewers. Special shoutout to our #ICCV2025 outstanding reviewers ๐Ÿซก iccv.thecvf.com/Conferences/...

๐Ÿš€ Just accepted to ICCV 2025! In DIY-SC, we improve foundational features using a light-weight adapter trained with carefully filtered and refined pseudo-labels. ๐Ÿ”ง Drop-in alternative to plain DINOv2 features! ๐Ÿ“ฆ Code + pre-trained weights available now. ๐Ÿ”ฅ Try it in your next vision project!

Olaf Dรผnkel@oduenkel.bsky.social ยท last yr.

Are you using DINOv2 for tasks that require semantic features? DIY-SC might be the alternative! It refines DINOv2 or SD+DINOv2 features and achieves a new SOTA on the semantic correspondence dataset SPair-71k when not relying on annotated keypoints! [1/6] genintel.github.io/DIY-SC

The CVML group at the @mpi-inf.mpg.de has been busy for CVPR. Check out our papers and come by the presentations!

Computer Vision and Machine Learning at MPI Informatics@cvml.mpi-inf.mpg.de ยท last yr.

๐ŸŽ‰ Exciting News #CVPR2025! Weโ€™re proud to announce that we have 5 papers accepted to the main conference and 7 papers accepted at various CVPR workshops this year! Weโ€™re looking forward to sharing our research with the community in Nashville! Stay tuned for more details! โ€ชโ€ช@mpi-inf.mpg.deโ€ฌ

Quantitative evaluation of diffusion model outputs is hard! We realized that we are often lacking metrics for comparing the quality of video and multi-view diffusion models. Especially the quantification of multi-view 3D consistency across frames is difficult. But not anymore: Introducing MET3R ๐Ÿงต

Bild