Dominik Schnaus

@schnaus.bsky.social

PhD student @ TUM with Daniel Cremers

New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

akoepke.github.io

Can you train a model for pose estimation directly on casual videos without supervision? Turns out you can! In our #CVPR2025 paper AnyCam, we directly train on YouTube videos and achieve SOTA results by using an uncertainty-based flow loss and monocular priors! ⬇️