Mickael Chen

@mickaelchen.bsky.social

Research Multimodal Generative AI, and now robotics. Generating MNIST digits for a decade.

DiffusionBlocks pub.sakana.ai/diffusionblo... from @sakanaai.bsky.social is one of these results that unlock a whole new branch of research papers. Implications of reframing as diffusion goes beyond memory efficiency. Test-time inpainting, guidance, and more would find new interpretations and uses.

DiffusionBlocks: Training Neural Networks One Block at a Time

A principled framework that converts a residual network into independently trainable blocks via a diffusion interpretation, achieving B× memory reduction without sacrificing performance.

pub.sakana.ai

Originally, the goal of generative models was just to capture the distribution of the dataset. Since then, we've shifted to maximize human preference and vibe checks. This must have introduced biases and made the models worse, at least on some aspect, at modelling the data distribution.

Kwang Moo Yi@kmyid.bsky.social · 6mo ago

Adamkiewicz et al., "When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators" Interesting that while we have "better" image generators, their usefulness as synthetic data generators is declining. Do we need a pivot?

🛠️ Already have a complex, pre-trained pipeline? If you are using bilinear interpolation anywhere, NAF acts as a strict drop-in replacement. Just swap it in. No retraining required. It’s literally free points for your metrics.📈

Bild

That was a cool project brillantly led by Ellington Kirby during his internship. We were curious if we could train diffusion models on sets of point coordinates. For images, this is a step towards spatial diffusion, with pixels reorganizing themselves, instead of diffusing in rgb values space only.

valeo.ai@valeoai.bsky.social · 9mo ago

LOGen: Toward Lidar Object Generation by Point Diffusion by: E. Kirby, @mickaelchen.bsky.social, R. Marlet, N. Samet tl;dr: a diffusion-based method producing lidar point clouds of dataset objects, with an extensive control of the generation 📄 arxiv.org/abs/2412.07385 Code: ✅

Wow, neet! Reannotation is key here. Conjecture: As we are get more and more well-aligned text-image data, it will become easier and easier to train models. This will allow us to explore both more streamlined and more exotic training recipes. More signals that exciting times are coming!

David Picard@davidpicard.eurosky.social · 2y ago

🚨 New preprint! How far can we go with ImageNet for Text-to-Image generation? w. @arrijitghosh.bsky.social @lucasdegeorge.bsky.social @nicolasdufour.bsky.social @vickykalogeiton.bsky.social TL;DR: Train a text-to-image model using 1000 less data in 200 GPU hrs! 📜https://arxiv.org/abs/2502.21318 🧵👇

🚗 Ever wondered if an AI model could learn to drive just by watching YouTube? 🎥👀 We trained a 1.2B parameter model on 1,800+ hours of raw driving videos. No labels. No maps. Just pure observation. And it works! 🤯 🧵👇 [1/10]

Bild

The plateau on training scaling and the shift to test-time scaling created favorable conditions for a competitor like DeepSeek to raise and catch up with OpenAI. Nah, I just made that up. Need to put more thoughts into this. 🤔

Better VQ-VAEs with this one weird rotation trick! I missed this when it came out, but I love papers like this: a simple change to an already powerful technique, that significantly improves results without introducing complexity or hyperparameters.

Bild
Post nicht verfügbar.

For AI to be fair and sustainable, we'd need to figure out attribution, i.e. "How much does training sample X contribute to model output Y?" Then the creator of sample X gets paid an amount proportional to what the user paid for the inference call that produced output Y.

A great place for students interested in AI/CV research internship. It's a very strong team, invested with all of their students. Check it out.

Andrei Bursuc@abursuc.bsky.social · 2y ago

🚨Hear! Hear! We have a few MSc research internship openings at valeo.ai in Paris for 2025 on computer vision & machine learning (yeah AI). You can find the openings in the link below along with the achievements of our amazing previous interns: valeoai.github.io/interns/ Join us!

ICYMI our PointBeV #CVPR2024 poster here's a quick talk by lead author Loïck Chambon. It brings a change of paradigm in multi-camera bird's-eye-view (BeV) segmentation via a flexible mechanism to produce sparse BeV points that can adapt to situation, task, compute www.linkedin.com/posts/andrei...

Andrei Bursuc on LinkedIn: #cvpr2024 #cvpr

In case you missed our PointBeV poster at #CVPR2024 here's a quick presentation by the lead author Loïck C.. PointBEV brings a change of paradigm in…

linkedin.com

The Cosmos suite of neural tokenizers for images & videos is impressive. Cosmos is trained on diverse high-res imgs & long-vids, scales well for both discrete & continuous tokens, generalizes to multiple domains (robotics, driving, egocentric ...) & has excellent runtime github.com/NVIDIA/Cosmo...

BildBildBildBild