The last project of my PhD is finally out! 🪴 It was a pleasure collaborating with Aimi on this work! We introduce A²BM: Alignment-Aware Bridge Matching, a new framework for image-to-image translation with weakly aligned image pairs. Paper 📄: arxiv.org/pdf/2607.16294
Nicolas Dufour
@nicolasdufour.bsky.social
Postdoc at Kyutai http://nicolas-dufour.github.io
I organize a 1 day workshop on Generative modelling @ENS Lyon, October 9th Call for oral/poster contributions is open; details at gdr-iasis.cnrs.fr/reunions/mod...
Modèles génératifs : diffusion, flow matching - GdR IASIS
Les demandes de prise en charge de missions par le GdR IASIS doivent parvenir à la gestionnaire du GdR avant le 25 septembre. Les modèles génératifs ont connu de récentes avancées spectaculaires, au p...
gdr-iasis.cnrs.fr
Excited to share that our paper on sprite-based image decomposition is accepted at TMLR! 🎉 Sprite-based models are highly interpretable but struggle to scale to complex, multi-object images. 1/3
There's a demo of the model btw: huggingface.co/spaces/nicol...
MIRO - a Hugging Face Space by nicolas-dufour
Multi-reward conditioned text-to-image diffusion (ICML 2026)
huggingface.co
I'm sadly not at ICML, but @arrijitghosh.bsky.social and @lucasdegeorge.bsky.social are presenting MIRO right now! Happening right now at poster board 2508!
I'm sadly not at ICML, but @arrijitghosh.bsky.social and @lucasdegeorge.bsky.social are presenting MIRO right now! Happening right now at poster board 2508!
1/ MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency. Wed, Jul 8, 2026 • 10:30 AM – 12:15 PM KST 📜 arxiv.org/abs/2510.25897 🖥️ nicolas-dufour.github.io/miro/ 🧬 huggingface.co/nicolas-dufo... With @arrijitghosh.bsky.social and @lucasdegeorge.bsky.social being on site
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typi...
arxiv.org
🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍
Dufour et al., "The FID Lottery: Quantifying Hidden Randomness in Generative Model Evaluation" We all know it's expensive to train multiple times, but we are now at a point where it is inevitable. Statistical significance should not be ignored. Don't bold over 1~2% differences.
I'll need to track how many time I refer to this paper. It's probably going to be my new language filler.
🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵
We explored the impact of variability sources in generative modeling. Turns out, we've been neglecting the error bars associated with training variability all along! We should aim to report results that we are sure of their scientific validity, instead of seed engineering!
🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵
Babe, stop everything! New favorite paper of the year is out! kyutai.org/fid-lottery/ arxiv.org/abs/2606.20536
The FID Lottery — Quantifying Hidden Randomness in Generative Model Evaluation
An interactive companion to 'The FID Lottery'. Every reported FID is the outcome of two lotteries — we measure how much they move the number.
kyutai.org
Hypnotizing to watch. Great work, co-authored by our very own @nicolasdufour.bsky.social
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/
𝗦𝘂𝗿𝗳𝗹𝗼: 𝗖𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝘁 𝟯𝗗 𝗦𝘂𝗿𝗳𝗮𝗰𝗲 𝗙𝗹𝗼𝘄 𝗠𝗼𝗱𝗲𝗹 𝘄𝗶𝘁𝗵 𝗚𝗹𝗼𝗯𝗮𝗹 𝗦𝘁𝗮𝘁𝗲 Antoine Guédon, Shu Nakamura, Nicolas Dufour ... Angjoo Kanazawa arxiv.org/abs/2606.13644 Trending on scholar-inbox.com
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/
We will release our code+data asap, please stay tuned! Thanks to my amazing coauthors: @nicolasdufour.bsky.social Shu Nakamura Jiahui Lei @kyotovision.bsky.social @akanazawa.bsky.social 📜arXiv: arxiv.org/abs/2606.13644 🔗Project: anttwo.github.io/surflo/ 💻Code (soon): github.com/Anttwo/Surflo 8/8
Surflo: Consistent 3D Surface Flow Model with Global State
Geometry is invariant to viewpoint, which makes any collection of images a redundant encoding of a single 3D state. Existing feed-forward reconstruction models fail to exploit this: per-view methods e...
arxiv.org
Check out our latest work! 🚀 We learn a global state and decode the point cloud pointwise, allowing to decode as many points as you want. Plus, we introduce some clever guidance tricks to ensure global consistency, yielding high-quality meshes from just a few views! 👇
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/
Surflo: Consistent 3D Surface Flow Model with Global State @antoine-guedon.bsky.social, Shu Nakamura, @nicolasdufour.bsky.social, Jiahui Lei, Ko Nishino, @akanazawa.bsky.social arxiv.org/abs/2606.13644
Playing with some post-processing on a small in-house-but-soon-to-be-released model.
The model is so fast and easy to use that I vibe-coded a small game with it in 1h 😅 Runs flawlessly on a consumer GPU if you're looking for a small local model to tinker with.
Everything is fully open-sourced, including the codebase, the model + all individual single reward model variants! 🌐 Site: nicolas-dufour.github.io/miro 📄 Paper: arxiv.org/abs/2510.25897 🛠️ Git: github.com/nicolas-dufo... 🤗 HF: huggingface.co/nicolas-dufo... 🎨 Demo: huggingface.co/spaces/nicol...
Open-source fueled the LLM revolution, but Physical AI hasn't fully benefited from this flywheel yet. Today, we're launching kesai.eu, our mission to democratize robotics research! First milestone: training a frontier-level self-driving policy using significantly less data than typically required.
KE:SAI — Open Science Autonomy Lab
KE:SAI is a Franco-German non-profit open science lab for scalable autonomous intelligence.
kesai.eu
Thrilled to share that MIRO is accepted to ICML 2026 @icmlconf.bsky.social ! 🎉 By training on the reward scores, we can simply condition the model on high rewards at inference time to guarantee top-tier, aligned outputs. We’ve updated our paper with some additional results!
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typi...
arxiv.org
We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉
We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉
👏 Folks! If you are curious about the Generative Modeling via Drifting paper, but you find it difficult to understand → I wrote a different interpretation of it. It's called: "An Expectation-Maximization interpretation of Generative Modeling via Drifting" davidpicard.github.io/pdf/An_Expec...
Excited to share my work as a Student Researcher at Google Zurich: UniGeoCLIP! 🌍🚀 W/ Eduard Trulls, Jan Hosang, @loicland.bsky.social & @pesarlin.bsky.social , we built a framework aligning 5 geospatial modalities in one space. Presented at EarthVision @ #CVPR2026. 🧵👇
New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
akoepke.github.io
Checkout our recent work, where we only need web images to learn a novel view generation model! We can navigate inside any image, without any video/multi view data or prior models! Congrats to Adrien for this great first PhD paper! (with @davidpicard.eurosky.social and @ptrkprz.bsky.social)
We're releasing OVIE, a novel view generation model trained entirely on single images. No multi-view datasets needed. Given a single image, it generates novel views of any scene in real time, running orders of magnitude faster than competing approaches.
I thought I would do a thread, but honestly the post is so good: kyutai.org/blog/2026-04... It explains "One View Is Enough! Monocular Training for In-the-Wild Novel View Generation" arxiv.org/abs/2603.23488 done in colab with the smart people at kyutai
OVIE: One View Is Enough!
Our mission is to build and democratize artificial general intelligence through open science.
kyutai.org
🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...
arxiv.org
🚨 Happy to announce CVPR@Paris'26 which will take place on June 1st in Paris. The goal of the event is to share a little bit of the conference before it happens. We will have poster sessions as well as several plenary talks by world-class speakers. info: cvprinparis.github.io/CVPR2026InPa...
CVPR@Paris 2026 June 1st
cvprinparis.github.io