Lucas Degeorge

@lucasdegeorge.bsky.social

PhD student at École Polytechnique (Vista) and École des Ponts (IMAGINE) Working on generative models

🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...

arxiv.org

🎉 Our work MIRO is accepted to #ICML2026 @icmlconf.bsky.social We integrate human preferences directly during pretraining with multi-reward conditioning. ⚡MIRO is 19x faster than baselines and 370x cheaper at inference! 🤗 Try out the models: huggingface.co/spaces/nicol... See you in Seoul 🇰🇷 !

MIRO - a Hugging Face Space by nicolas-dufour

Multi-reward conditioned text-to-image diffusion (ICML 2026)

huggingface.co

Nicolas Dufour@nicolasdufour.bsky.social · 4mo ago

Thrilled to share that MIRO is accepted to ICML 2026 @icmlconf.bsky.social ! 🎉 By training on the reward scores, we can simply condition the model on high rewards at inference time to guarantee top-tier, aligned outputs. We’ve updated our paper with some additional results!

We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉

Bild

Final note: I'm (we're) tempted to organize a challenge on that topic as a workshop at a CV conf. ImageNet is the only source of images allowed and then you compete to get the bold numbers. Do you think there would be people in for that? Do you think it would make for a nice competition?

Text-to-image models are trained on billions of data. But, is it necessary? Our "How far can we go with ImageNet for T2I generation?‬" @lucasdegeorge.bsky.social @arrijitghosh.bsky.social @nicolasdufour.bsky.social @davidpicard.bsky.social shows that no, if we are careful arxiv.org/abs/2502.21318

Bild
David Picard@davidpicard.eurosky.social · 2y ago

🚨 New preprint! How far can we go with ImageNet for Text-to-Image generation? w. @arrijitghosh.bsky.social @lucasdegeorge.bsky.social @nicolasdufour.bsky.social @vickykalogeiton.bsky.social TL;DR: Train a text-to-image model using 1000 less data in 200 GPU hrs! 📜https://arxiv.org/abs/2502.21318 🧵👇

Wow, neet! Reannotation is key here. Conjecture: As we are get more and more well-aligned text-image data, it will become easier and easier to train models. This will allow us to explore both more streamlined and more exotic training recipes. More signals that exciting times are coming!

David Picard@davidpicard.eurosky.social · 2y ago

🚨 New preprint! How far can we go with ImageNet for Text-to-Image generation? w. @arrijitghosh.bsky.social @lucasdegeorge.bsky.social @nicolasdufour.bsky.social @vickykalogeiton.bsky.social TL;DR: Train a text-to-image model using 1000 less data in 200 GPU hrs! 📜https://arxiv.org/abs/2502.21318 🧵👇