Philippe Weinzaepfel

@weinzaepfelp.bsky.social

Principal Research Scientist in Computer Vision at Naver Labs Europe https://philippeweinzaepfel.github.io/

#ECCV2026 paper: A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world 966 *REAL* nav episodes (!!) performed by S. Janny with Dino-v3, Dino-v2, DUNE, VC1, AM-RADIO encoders show that patch features can be bottlenecked to 1 value ➡️ affordances emerge. 1/8

Bild

For your Embodied AI task you want a recurrent model with constant complexity per step, but you don't want to lose the power of transformers (which store the full obs history and attend to it)? Do not despair, we have your back. We distill transformers into recurrent transformers 1/8

Bild

Interested in designing the next generation of FF3D reconstruction models for real-life use cases? The Geometric Deep Learning team at the root of the foundational research lines of CroCo and DUSt3R is looking for a research scientist to help us!

🧍‍♀️ Introducing Anny: an open, interpretable, and differentiable human body model for all ages. Grounded in anthropometric data (MakeHuman) & WHO stats, Anny offers: 🧠 Interpretable shape control 👶👩‍🦳 Unified from infants to elders 🧩 Versatile for fitting, synthesis & HMR 🌍 Open under Apache 2.0

Romain Brégier@rbregier.bsky.social · 10mo ago

Meet Anny, our Free (Apache 2.0) and Interpretable Human Body Model for all ages. Anny is built upon #MakeHuman and enables achieving SOTA performance in Human Mesh Recovery. ArXiv: arxiv.org/abs/2511.03589 Demo: anny-demo.europe.naverlabs.com Code: github.com/naver/anny

"Sliding is all you need" (aka "What really matters in image goal navigation") has been accepted to 3DV 2026 (@3dvconf.bsky.social) as an Oral presentation! By Gianluca Monaci, @weinzaepfelp.bsky.social and myself. @naverlabseurope.bsky.social

Christian Wolf@chriswolfvision.bsky.social · last yr.

In a new paper led by Gianluca Monaci, with @weinzaepfelp.bsky.social and myself, we explore the relationship between rel pose estimation and image goal navigation and study different architectures: late fusion, channel cat (w/ or w/o space2depth) and cross-attention. arxiv.org/abs/2507.01667 🧵1/5