Antoine Guédon

@antoine-guedon.bsky.social

Postdoctoral researcher in computer vision at Ecole polytechnique. I'm interested in 3D Reconstruction, Radiance Fields, Gaussian splatting, 3D Scene Rendering, 3D Scene Understanding, etc. Webpage: https://anttwo.github.io/

🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍

Bild

We explored the impact of variability sources in generative modeling. Turns out, we've been neglecting the error bars associated with training variability all along! We should aim to report results that we are sure of their scientific validity, instead of seed engineering!

Kyutai@kyutai-labs.bsky.social · 2mo ago

🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵

🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵

What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/

🔴FROM BLOBS TO SPOKES🚲 We released paper and code for GaussianWrapping, our latest work on RGB-to-mesh! We introduce explicit geometric field formulas for Gaussians (occupancy&normals), allowing for fast and sharp surface reco (see bicycle spokes). So happy about this work!🤩

Diego Gomez@diegoxgomez.bsky.social · 4mo ago

1/n 🧵 Introducing Gaussian Wrapping — a principled framework for extracting high-quality meshes from 3DGS! 🚲 We recover thin structures, like bicycle spokes, where all prior methods fail. Follow the thread for a brief overview and links!

🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...

arxiv.org

We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉

Bild

1/n🚀Gaussians > Differentiable function > Mesh? Check out our new work: MILo: Mesh-In-the-Loop Gaussian Splatting! 🎉Accepted to SIGGRAPH Asia 2025 (TOG) MILo is a novel differentiable framework that extracts meshes directly from Gaussian parameters during training. 🧵👇

I'm at #CVPR2025 to present our paper 🍵MAtCha Gaussians🍵, today Friday afternoon, Hall D, Poster 53! If you're in Nashville and want to discuss detailed 3D mesh reconstruction from sparse or dense RGB images, let's connect! @kyotovision.bsky.social

Antoine Guédon@antoine-guedon.bsky.social · last yr.

💻We've released the code for our #CVPR2025 paper MAtCha! 🍵MAtCha reconstructs sharp, accurate and scalable meshes of both foreground AND background from just a few unposed images (eg 3 to 10 images)... ...While also working with dense-view datasets (hundreds of images)!

💻We've released the code for our #CVPR2025 paper MAtCha! 🍵MAtCha reconstructs sharp, accurate and scalable meshes of both foreground AND background from just a few unposed images (eg 3 to 10 images)... ...While also working with dense-view datasets (hundreds of images)!

BildBildBildBild

🤔 What if embedding multimodal EO data was as easy as using a ResNet on images? Introducing AnySat: one model for any resolution (0.2m–250m), scale (0.3–2600 hectares), and modalities (choose from 11 sensors & time series)! Try it with just a few lines of code:

Bild