Benjamin Lefaudeux 🇺🇦

@bentheegg.bsky.social

Back to France after some time in sunny California and happy Copenhagen. Mistral, Photoroom, Meta (xformers, FairScale, R&D), EyeTribe (acq) Mostly writing around AI

In 2012 when I had to clean data it seemed natural to look for rules I could use to clean it. Now it seems natural to model the noise, find new clean data it can destroy, and then train a model to reverse the process. Machine learning makes you a sicko.

Three things to note about this: 1) AI has obvious utility to many, this is a tremendous amount of use already 2) There is room for multiple frontier model providers, at least for now 3) Any losses from subsidizing cost of AI use (and it is not clear this is happening) are now relatively small

Bild

1/ Can open-data models beat DINOv2? Today we release Franca, a fully open-sourced vision foundation model. Franca with ViT-G backbone matches (and often beats) proprietary models like SigLIPv2, CLIP, DINOv2 on various benchmarks setting a new standard for open-source research.

Bild

In the coming age of agents, I think vibe coding will die out, same lasting power as prompt engineering. For things LLMs excell at, you might as well stick to higher level directives and let it own the work, Claude Code is a good example. 1/2

this is probably why Meta was able to poach OpenAI ppl aside from the absolute piles of cash, Sama is very SV-minded and can’t imagine building apart from a product a lot of accelerationists see things differently, more broadly, and ids dissatisfying to be forced into a product box

mr. TIM@timkellogg.me · last yr.

explaining why they open sourced — to ensure that it’s broadly useful OpenAI self-admits that they optimize their models for ChatGPT, o3 was made for DeepResearch Moonshot was dissatisfied with that

For a closed-source ChatBot service, users have no idea what workflow or how many models are behind it. I've heard rumors that some major companies have dozens of models, hundreds of scenario classifications, and countless workflows behind their interfaces, claiming this is an "MoE model." Under "application-first" or "user experience-first" values, this approach is very natural and far more cost-effective than a single model.
But this clearly isn't what AGI should look like. For a startup like Kimi, this approach not only makes you increasingly mediocre and greatly hinders technical progress but also makes it impossible to compete with major companies that have PMs polishing every button.

Little bit of personal news, shared in other circles already: I'm moving to Mistral in August, after three years at Photoroom. I'm really proud of what we built in the ML team with relatively limited means, lasting SOTA on the existing foundations (saliency segmentation) while growing a lot on genAI

Still haven't tried Cursor, but I recently moved from Github Copilot to Continue with Codestral (free API), and it's absurd how much better Continue with Codestral is (vs. Copilot with expensive and slow models). Made me realize that there is zero moat in this field, at least for Copilot.

datago now available with webdataset compatibility (streaming tarballs, so you get the data as it arrives). Just pip install datago and give it a whirl if you'd like ? Speed without the dataloader processes, and typical ViT/DiT pre-processing baked in. example code here github.com/Photoroom/da...

datago/python/benchmark_webdataset.py at main · Photoroom/datago

A Rust-based data loader which can be used from Python. Processing data per sample at GB/s speeds, covering various use cases eventually. - Photoroom/datago

github.com

SageAttention3 paper reads great, and looks like B200s just got a good value boost. QAT or PTQ-free use of FP4, I expected this to be much more complicated or come later to be honest. Only at the attention level and LLMs are most often MLP bottlenecked but stil arxiv.org/abs/2505.11594

SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training

The efficiency of attention is important due to its quadratic time complexity. We enhance the efficiency of attention through two key contributions: First, we leverage the new FP4 Tensor Cores in Blac...

arxiv.org

Big Marigold update! Last year, we showed how to turn Stable Diffusion 2 into a SOTA depth estimator with a few synthetic samples and 2–3 days on just 1 GPU. Today's release features: 🏎️ 1-step inference 🔢 New modalities 🫣 High resolution 🧨 Diffusers support 🕹️ New demos 🧶👇

Striking in retrospect how some leaders position at the time of the first atomic bomb (“other countries won’t get it”), Truman for instance, then space age, then now AI age, rhyme. Same as before, I think a bunch of places are bound to be SOTA AI, ideas and progress cannot be pinned to a wall