Probably the best model report this year: multiple infrastructure pieces cohering to fix computation and network to hold 3T parameters.
You’ve waited long enough, the Kimi K3 open weights & full tech report are here! github.com/MoonshotAI/K...
Probably the best model report this year: multiple infrastructure pieces cohering to fix computation and network to hold 3T parameters.
You’ve waited long enough, the Kimi K3 open weights & full tech report are here! github.com/MoonshotAI/K...
My one issue with Nolan’s version so far: it cuts everything that makes Odyssey a self-aware proto-philosophical text, especially Scheria which has already all features of a Plato’s myth. One of the major pre-Socratic work, Parmenides’ poem, is literally an Odyssey rewrite.
And new technical blogpost by Pleias application team on deploying small reasoning models for edge devices : featuring cache context management on Rasperry, designing system orchestration under constraints (reranker, chunking) and model specialization. pleias.ai/blog/local-a...
Alors deux minutes d’explications : 1. Les modèles à poids ouverts sont protégés par la norme safetensors : par définition ils n’embarquent pas de code. Jamais vu un cas de backdoor.
IA : le test d’un modèle chinois à la direction générale du Trésor interrompu à cause de « réponses orientées » ou « biaisées »
Announcing the first industrial application of SYNTH: we trained a 600m reasoning model for one of the largest infrastructure in the world, the subway of Paris. pleias.ai/blog/sillon-...
Doing the most responsible thing an European AI labs can do after this weekend: shipping a blogpost. Why the EU can't into AI, how it's not about compute, but actual skill issue and failing for years to build an actual training ecosystem. pleias.ai/blog/fable-eu
Ok if the only answer to all this is billions of Mistral subsidies, I’m out.
Not having any EU competitive labs (plural intended) is about to get very painful.
Not having any EU competitive labs (plural intended) is about to get very painful.
Anthropic says it is disabling Fable 5 and Mythos 5 for all customers after the US government issued an export control order, citing national security concerns (Anthropic) Main Link | Techmeme Permalink
After months of delay, here comes the successor post to "The model is the product": the AI decoupling. All about MoE high margin economics, synthetic pretraining weakening commoditization and the new push toward Model IP. vintagedata.org/blog/posts/t...
So more corporate news: we're launching an early beta access to the synth pipelines that originally created SYNTH and have been further refined and enhanced through the last few months.
And new Pleias release in partnership with GSMA: CommonLingua, a 2.35M parameters model for language detection, currently performing best on the CommonLID benchmark by a wide margin. huggingface.co/PleIAs/Commo...
Currently presenting the Common Corpus poster at #ICLR2026 Pavilion 3 (spot 1614) with Pavel Chizhov, if you want to come and say hi.
First day of #ICLR2026 posters: graph is one of the most interesting area right now. Original methods involving bold architecture choices (beyond classic GNN) and I do see an immediate use for it in synthetic/agentic pipelines as we need to perform search at increasing depth levels.
At least for now Brazil is 10/10 on things that matter (breakfast)
More formal announcement: I'll be at ICLR this week to present the oral on Common Corpus with Pavel Chizhov. Happy to discuss anything data research, synth pretraining, models and products.
Parameter ceiling seems ridiculously low now that we can generate data at will. For a complex token classification task (editorial structure detection) i'm maxing out at 4M.
I doubt anyone ever did that, but there is a philosophical thesis to write on the ties between Homotopy Type Theory and ancient logic.
Don't know how to say to Mistral that some of it already exists…
Just realized Common Corpus is (again!) in Anthropic's Transformers Circuit — new chapter on emotions transformer-circuits.pub/2026/emotion...
Hum. I have a slight Chinese doubt over Meta’s new model…
Meanwhile, just occurred to me that the "AI con" is still on a book tour.
circuit transformers once more confirmed confirmed to be the one oblique access to anthropic core research: most interesting parts of the 240 pages mythos report.
Maybe it’s too soon to talk about a vibe shift toward data research in AI, but definitely easier to get paper accepted.
One of the weirdest part of my job right now is being naturally led to read philosophical works for practical data pipeline design.
Looks like I’m bound to reinvent all pretraining tooling around sub-4M bytes model. Classification done, language detection next.
I like how Anthropic is just obliquely releasing their work on recursive self-improvement.
on a different note: happily featured in the main french ranking for young economic leaders