Guillaume Astruc

@gastruc.bsky.social

2nd Year PhD Student from Imagine-ENPC/IGN/CNES Working on Self-supervised Cross-modal Geospatial Learning. Personal WebPage: https://gastruc.github.io/

🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍

Bild

What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/

🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...

arxiv.org

We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉

Bild

🛰️ At #CVPR2025 presenting "AnySat: An Earth Observation Model for Any Resolutions, Scales, and Modalities" - Saturday afternoon, Poster 355! If you're here and want to discuss geolocation or geospatial foundation models, let's connect!

I will be presenting our work on the detection of archaeological looting with satellite image time series at CVPR 2025 EarthVision workshop tomorrow! Honored and grateful that this paper received the best student paper award!

Bild

📢 New preprint! “When majority rules, minority loses: bias amplification of gradient descent” We often blame biased data but training also amplifies biases. Our paper explores how ML algorithms favor stereotypes at the expense of minority groups. ➡️ arxiv.org/abs/2505.13122 (1/3)

When majority rules, minority loses: bias amplification of gradient descent

Despite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, ...

arxiv.org

We've added new experiments demonstrating robust generalization capabilities! Notably, AnySat shows strong performance on HLS Burn Scars - a sensor never seen during pretraining! 🔥🛰️ Check it out: 📄 Paper: arxiv.org/abs/2412.14123 🌐 Project: gastruc.github.io/anysat

Bild
Loïc Landrieu @loicland.bsky.social · last yr.

AnySat has been accepted as a ✨ highlight at #CVPR2025! See you in Nashville 🎉 We’ll also be presenting this work at: 📍 @egu.eu on 02/04 in Vienna 📍 @esa.int / NASA Workshop on Foundation Models on 05/04 in Rome

💻We've released the code for our #CVPR2025 paper MAtCha! 🍵MAtCha reconstructs sharp, accurate and scalable meshes of both foreground AND background from just a few unposed images (eg 3 to 10 images)... ...While also working with dense-view datasets (hundreds of images)!

BildBildBildBild

Weights for CAD are finally available. It's one of the smallest diffusion models on the market, achieving performance close to SD and Pixart, featuring a Perceiver-like architecture. We leverage our coherence aware training to improve the textual understanding

David Picard@davidpicard.eurosky.social · last yr.

🚨 Just a quick note that following requests, we trained a 512px version of our Coherence-Aware Diffusion model (CVPR'24) and updated the paper on arxiv: arxiv.org/abs/2405.20324 It has a package and pretrained models! 🖥️ nicolas-dufour.github.io/cad.html 🤖 github.com/nicolas-dufo...

🤔 What if embedding multimodal EO data was as easy as using a ResNet on images? Introducing AnySat: one model for any resolution (0.2m–250m), scale (0.3–2600 hectares), and modalities (choose from 11 sensors & time series)! Try it with just a few lines of code:

Bild