🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍
Antoine Guédon
@antoine-guedon.bsky.social
Postdoctoral researcher in computer vision at Ecole polytechnique. I'm interested in 3D Reconstruction, Radiance Fields, Gaussian splatting, 3D Scene Rendering, 3D Scene Understanding, etc. Webpage: https://anttwo.github.io/
We explored the impact of variability sources in generative modeling. Turns out, we've been neglecting the error bars associated with training variability all along! We should aim to report results that we are sure of their scientific validity, instead of seed engineering!
🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵
🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵
𝗦𝘂𝗿𝗳𝗹𝗼: 𝗖𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝘁 𝟯𝗗 𝗦𝘂𝗿𝗳𝗮𝗰𝗲 𝗙𝗹𝗼𝘄 𝗠𝗼𝗱𝗲𝗹 𝘄𝗶𝘁𝗵 𝗚𝗹𝗼𝗯𝗮𝗹 𝗦𝘁𝗮𝘁𝗲 Antoine Guédon, Shu Nakamura, Nicolas Dufour ... Angjoo Kanazawa arxiv.org/abs/2606.13644 Trending on scholar-inbox.com
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/
Surflo: Consistent 3D Surface Flow Model with Global State @antoine-guedon.bsky.social, Shu Nakamura, @nicolasdufour.bsky.social, Jiahui Lei, Ko Nishino, @akanazawa.bsky.social arxiv.org/abs/2606.13644
CVPR@Paris 2026 🇫🇷 — June 1st, co-organised by ELLIS Unit Paris. A one-day local event ahead of CVPR, open to all. Oral & poster sessions for CVPR 2026, CVPR workshops & ICLR 2026 papers. 🔗 https://cvprinparis.github.io/CVPR2026InParis/
Thrilled to share that MIRO is accepted to ICML 2026 @icmlconf.bsky.social ! 🎉 By training on the reward scores, we can simply condition the model on high rewards at inference time to guarantee top-tier, aligned outputs. We’ve updated our paper with some additional results!
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typi...
arxiv.org
We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉
🔴FROM BLOBS TO SPOKES🚲 We released paper and code for GaussianWrapping, our latest work on RGB-to-mesh! We introduce explicit geometric field formulas for Gaussians (occupancy&normals), allowing for fast and sharp surface reco (see bicycle spokes). So happy about this work!🤩
1/n 🧵 Introducing Gaussian Wrapping — a principled framework for extracting high-quality meshes from 3DGS! 🚲 We recover thin structures, like bicycle spokes, where all prior methods fail. Follow the thread for a brief overview and links!
🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...
arxiv.org
I’ll be at #SIGGRAPHAsia2025 next week presenting our paper MILo! Join the Neural Fields and Surface Reconstruction session on Tuesday, December 16. If you’ll be in Hong Kong and would like to discuss research, or grab a coffee ☕️ feel free to reach out.
1/n🚀Gaussians > Differentiable function > Mesh? Check out our new work: MILo: Mesh-In-the-Loop Gaussian Splatting! 🎉Accepted to SIGGRAPH Asia 2025 (TOG) MILo is a novel differentiable framework that extracts meshes directly from Gaussian parameters during training. 🧵👇
We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉
Familiar names among #ICCV2025 Outstanding Reviewers from our team 😇 Antoine Guédon @antoine-guedon.bsky.social Sinisa Stekovic Renaud Marlet 👏 @iccv.bsky.social iccv.thecvf.com/Conferences/...
2025 ICCV Program Committee
iccv.thecvf.com
1/n🚀Gaussians > Differentiable function > Mesh? Check out our new work: MILo: Mesh-In-the-Loop Gaussian Splatting! 🎉Accepted to SIGGRAPH Asia 2025 (TOG) MILo is a novel differentiable framework that extracts meshes directly from Gaussian parameters during training. 🧵👇
I'm at #CVPR2025 to present our paper 🍵MAtCha Gaussians🍵, today Friday afternoon, Hall D, Poster 53! If you're in Nashville and want to discuss detailed 3D mesh reconstruction from sparse or dense RGB images, let's connect! @kyotovision.bsky.social
💻We've released the code for our #CVPR2025 paper MAtCha! 🍵MAtCha reconstructs sharp, accurate and scalable meshes of both foreground AND background from just a few unposed images (eg 3 to 10 images)... ...While also working with dense-view datasets (hundreds of images)!
Behind every great conference is a team of dedicated reviewers. Congratulations to this year’s #CVPR2025 Outstanding Reviewers! cvpr.thecvf.com/Conferences/...
#CVPR2025 Fri June 13 (PM) ✨ Highlight 🍵 MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views @antoine-guedon.bsky.social @kyotovision.bsky.social 📄 pdf: arxiv.org/abs/2412.06767 🌐 webpage: anttwo.github.io/matcha/
💻We've released the code for our #CVPR2025 paper MAtCha! 🍵MAtCha reconstructs sharp, accurate and scalable meshes of both foreground AND background from just a few unposed images (eg 3 to 10 images)... ...While also working with dense-view datasets (hundreds of images)!
🔥🔥🔥 CV Folks, I have some news! We're organizing a 1-day meeting in center Paris on June 6th before CVPR called CVPR@Paris (similar as NeurIPS@Paris) 🥐🍾🥖🍷 Registration is open (it's free) with priority given to authors of accepted papers: cvprinparis.github.io/CVPR2025InPa... Big 🧵👇 with details!
Starter pack including some of the lab members: go.bsky.app/QK8j87w
1/13 🐊 Introducing our latest work on improving relative camera pose regression with a novel pre-training approach Alligat0R (arxiv.org/abs/2503.07561)! @gbourmaud.bsky.social @vincentlepetit.bsky.social
🤔 What if embedding multimodal EO data was as easy as using a ResNet on images? Introducing AnySat: one model for any resolution (0.2m–250m), scale (0.3–2600 hectares), and modalities (choose from 11 sensors & time series)! Try it with just a few lines of code:
⚠️Reconstructing sharp 3D meshes from a few unposed images is a hard and ambiguous problem. ☑️With MAtCha, we leverage a pretrained depth model to recover sharp meshes from sparse views including both foreground and background, within mins!🧵 🌐Webpage: anttwo.github.io/matcha/
MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views @antoine-guedon.bsky.social, Tomoki Ichikawa, Kohei Yamashita, Ko Nishino tl;dr: underlying scene geometry mesh->an Atlas of Charts->render with 2D Gaussian surfels arxiv.org/abs/2412.06767
Welcome to all the newcomers! Here's a starter pack for the 3D Computer Vision community: go.bsky.app/Cfm9XFe