lebellig

@lebellig.bsky.social

Postdoc @INRIA, Ockham team, on generative models. Previously intern @SonyCSL, @Ircam, @INRIA 🌎 Personal website: https://lebellig.github.io/

Here's the tale of how @jder.bsky.social and I scaled Samudra, a neural ocean emulator capable of predicting 8 years of the ocean on a single GPU, to operate at a full 1/4° resolution (16x the size in bytes). It was quite a humbling process.

Open Athena@openathena.ai · 6d ago

Simulating ocean climate takes a supercomputer 4,600+ CPU cores to produce 12 simulated years per day (SYPD). Samudra 2 produces 4,800 SYPD on 1 GPU at the same resolution. In a new blog, @al.merose.com reports on Samudra, a neural ocean emulator built in collaboration with NYU & MIT: bit.ly/oa-ss

NeurIPS submissions confirmed to be a heat-loving species. Warmer year, bigger bloom. Every degree we add, the deadline gets denser 🌻 Good news for the field, we're having a really good growing season 👨‍🌾

Bild

The last project of my PhD is finally out! 🪴 It was a pleasure collaborating with Aimi on this work! We introduce A²BM: Alignment-Aware Bridge Matching, a new framework for image-to-image translation with weakly aligned image pairs. Paper 📄: arxiv.org/pdf/2607.16294

BildBild

I'd like to announce that at @openathena.ai, @jder.bsky.social and I helped @m2lines.bsky.social release Samudra 2. We scaled this neural ocean emulator to train on 16x the size of data in bytes on the same hardware budget. We can now skillfully predict 8 years of the ocean on a single GPU at a 1/4°

🌊 Samudra 2: A Fast, Cheap AI Ocean Model, Now at the Scale That Matters

M²LInES’ neural ocean emulator now runs multi-year simulations at eddy-permitting resolution on a single GPU, turning a supercomputer-scale…

medium.com

🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍

Bild

We explored the impact of variability sources in generative modeling. Turns out, we've been neglecting the error bars associated with training variability all along! We should aim to report results that we are sure of their scientific validity, instead of seed engineering!

Kyutai@kyutai-labs.bsky.social · 2mo ago

🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵

🎆 New paper! "Random Process Flow Matching: Generative Implicit Representations of Multivariate Random Fields", by Julien Lalanne, accepted to ICML'26 🥳 We're proposing flow-matching for inpainting in ultra-sparse setup, with applications to seismic interpolation. 📜 arxiv.org/abs/2605.28625 1/

Random Process Flow Matching: Generative Implicit Representations of Multivariate Random Fields

Generative modeling provides a powerful framework for learning data distributions. These models initially relied on probabilistic methods such as Gaussian Processes (GP) for uncertainty-aware predicti...

arxiv.org

"Accept (spotlight)" at ICML'26 😎 Our paper brings particle filters back to life: autoregressive diffusion models + posterior sampling yield optimal proposals for Bayesian filtering, scaling up to GenCast-sized systems. arxiv.org/abs/2605.20028 w/ Thomas Savary and @francois-rozet.bsky.social

Training-Free Bayesian Filtering with Generative Emulators

Bayesian filtering is a well-known problem that aims to estimate plausible states of a dynamical system from observations. Among existing approaches to solve this problem, particle filters are theoret...

arxiv.org

ICML Conference@icmlconf.bsky.social · 3mo ago

Congrats again to authors of accepted #ICML2026 papers! The camera-ready deadline is 5/28. Drawing your attention to two specific features: 1. As last year, to help communicate research to a broad audience, papers will have lay summaries. Tips & details in blog 1/3

My PhD thesis manuscript will be available in the coming months, but I’ve written two blog posts based on the related work chapter: 1. Generative modeling with flow-based models 🪚 2. Data-translation with flow and diffusion bridges 🔨 Open to feedback and discussions! lebellig.github.io/blog/

BildBildBild

My PhD thesis manuscript will be available in the coming months, but I’ve written two blog posts based on the related work chapter: 1. Generative modeling with flow-based models 🪚 2. Data-translation with flow and diffusion bridges 🔨 Open to feedback and discussions! lebellig.github.io/blog/

BildBildBild

Delighted to have successfully defended my PhD thesis on "Generative models for Earth Observation, from denoising to domain adaptation" 🪴 This wouldn’t have been possible without the support of my colleagues, my PhD advisor, and my family and friends. Thank you all!

BildBildBild

Two research blog posts that look really interesting 👀 "Teaching AI to Invent Enzymes Nature Never Imagined", DISCO: Diffusion for Sequence-structure CO-design disco-design.github.io "How to Generate Text in One Step", Flow Map Language Models one-step-lm.github.io/blog/

DISCO — Teaching AI to Invent Enzymes Nature Never Imagined

DISCO is a multimodal generative model that co-designs protein sequence and 3D structure to create entirely new enzymes for reactions never seen in biology.

disco-design.github.io

🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...

arxiv.org

For those interested in normalized gradient methods and optimal transport: I introduce a new class of "spectral" Wasserstein distances for which spectrally normalized gradient descent (Muon but without momentum and small step size ...) is a spectral-W gradient flow: arxiv.org/abs/2604.04891

Muon Dynamics as a Spectral Wasserstein Flow

Gradient normalization is central in deep-learning optimization because it stabilizes training and reduces sensitivity to scale. For deep architectures, parameters are naturally grouped into matrices ...

arxiv.org