Nicolas Dufour

@nicolasdufour.bsky.social

Postdoc at Kyutai http://nicolas-dufour.github.io

The last project of my PhD is finally out! 🪴 It was a pleasure collaborating with Aimi on this work! We introduce A²BM: Alignment-Aware Bridge Matching, a new framework for image-to-image translation with weakly aligned image pairs. Paper 📄: arxiv.org/pdf/2607.16294

BildBild

Excited to share that our paper on sprite-based image decomposition is accepted at TMLR! 🎉 Sprite-based models are highly interpretable but struggle to scale to complex, multi-object images. 
 1/3

Bild

🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍

Bild

Dufour et al., "The FID Lottery: Quantifying Hidden Randomness in Generative Model Evaluation" We all know it's expensive to train multiple times, but we are now at a point where it is inevitable. Statistical significance should not be ignored. Don't bold over 1~2% differences.

Bild

We explored the impact of variability sources in generative modeling. Turns out, we've been neglecting the error bars associated with training variability all along! We should aim to report results that we are sure of their scientific validity, instead of seed engineering!

Kyutai@kyutai-labs.bsky.social · 2mo ago

🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵

What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/

Check out our latest work! 🚀 We learn a global state and decode the point cloud pointwise, allowing to decode as many points as you want. Plus, we introduce some clever guidance tricks to ensure global consistency, yielding high-quality meshes from just a few views! 👇

Antoine Guédon@antoine-guedon.bsky.social · 2mo ago

What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/

We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉

Bild

New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

akoepke.github.io

Checkout our recent work, where we only need web images to learn a novel view generation model! We can navigate inside any image, without any video/multi view data or prior models! Congrats to Adrien for this great first PhD paper! (with @davidpicard.eurosky.social and @ptrkprz.bsky.social)

Kyutai@kyutai-labs.bsky.social · 4mo ago

We're releasing OVIE, a novel view generation model trained entirely on single images. No multi-view datasets needed. Given a single image, it generates novel views of any scene in real time, running orders of magnitude faster than competing approaches.

🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...

arxiv.org