If you're teaching a class on post-training (or part of a course) and my book, slides, code or videos don't help you, please lmk how I can improve it! Has been a ton of work getting everything done and part of the ROI is hope that it helps more education work grows around LLMs.
Nathan Lambert
@natolambert.bsky.social
A LLN - large language Nathan - (RL, RLHF, society, robotics), athlete, yogi, chef Writes http://interconnects.ai Prev Ai2/Olmo, HuggingFace, Berkeley, and normal places
Another quick q&a video as I wrap up the course soon. These ended up all being on various details in getting a post-training recipe right. www.youtube.com/watch?v=6StK...
Post-Training Recipe Questions Answered | Q&A 3, Post-Training Course (w/ The RLHF Book)
In this Q&A I answer 6 questions and post-training recipes, how to think about trade-offs, and how the different loss functions work. Definitely a jam-packed and bit random, but fun nonetheless!…
youtube.com
Introducing our Artifacts Hub and Adoption Dashboard Scaling our curation and measurement of the open ecosystem as we feel the acceleration of releases. Artifacts Hub: artifactshub.ai Adoption Dashboard: dashboard.interconnects.ai Explanation: www.interconnects.ai/p/introducin...
Artifacts 23: Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier. In this issue, a total of 24 models from june/july you should be aware of. www.interconnects.ai/p/latest-ope...
The pace of progress on models from so many organizations at once is genuinely incredible. Building LLMs isn't driven by rare secrets, but consistent effort, mass capital, and effective organization design. It is great that know-how of such a powerful technology is diffused.
Another new lecture! Lecture 11 is a tool-use/function calling/agentic 101. I almost skipped this one, as this chapter started as the only skill-specific topic in the book, but since writing it tool-use has only become more foundational to modern models.
Frontier labs will be viable businesses by being able to integrate and optimize inference at lower cost/performance than most other models. They'll have a margin advantage on open models for the foreseeable future. This is one of those flex's imo and I doubt they're losing money.
New podcast/lecture combo -- a case study in the messy details of Olmo 3 post training & DPO with Scott Geng. It's rare to make time for these discussions. www.youtube.com/watch?v=rhA7...
From Academic Research to a Frontier LLM: A Case Study in DPO | RLHF Book Course, Conversation 2
The second researcher conversation of the course! We discuss what it takes to land a well-grounded academic result into a near-frontier model. Scott Geng and I worked together on DPO for Olmo 3, and…
youtube.com
Lecture 10 of my course! Nominally on regularization in RL, so I discuss the evolving role of the KL penalty in RL, but also a set of nice RL papers that explain what RL helps models generalize better than SFT -- with theory supporting it.
Kimi K3 with more likes than downloads on HuggingFace is definitely showing us a glimpse of the future on open models. It's way less about individual access, and more of a distributed platform layer for companies.
Making talks with AI agents is awesome. I just told Fable to make a slide with real data on the KL distance from one of our reference Olmo 2 models and it made this with the wandb api (I edited text slightly).
New (shorter) lecture! Over-optimization, foundations of reward hacking, sycophancy, verbosity, etc. In recording this, I realized that rubrics are going to be prone to overopt in a way like reward models, where RLVR is its own thing. Fundamentals, history, and reflections! youtu.be/y04JhXpiI4s
Over-Optimization and RLHF’s Bad Reputation | Post-Training Course, Lecture 9
In this lecture, we discuss how RLHF got a bad reputation, how over-optimization compares to over-fitting, and the nuance in getting preference tuning right. It's a shorter one, in order to keep…
youtube.com
A big day for open models, I feel like I can relax a little, soak it in. An incredibly powerful coalition said "we believe, we care." That's rare on any topic.
Insane numbers for opus 5, the power of faster iteration speed + scaled RL (Fable too big to RL as well, yet). And on safeguards "Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5.". www.anthropic.com/news/claude-...
The real comparison to Moore's law for AI isn't scaling laws, but rather the intelligence efficiency that we gain year-over-year using models to get better at training models.
If you're looking for the latest adoption data on open models in US v China v globally, we built a small dashboard with the big picture and per-org numbers. US's role is slowly growing, but still way behind China/Qwen. dashboard.interconnects.ai
At night I dream of a distillation debate grounded in public, technical info, not reading tea leaves of backroom deals and political intrigue. Then I wake up and I'm crushed by reality of chaos, potentially classified information, and a spiraling global AI ecosystem.
My book is the number 1 AI bestseller on Amazon. Is a success, even if that just lasts for a day :)! Thanks all for the support. rlhfbook.com
Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next A podcast with Florian Brand. YouTube: www.youtube.com/watch?v=XsBy... Interconnects: www.interconnects.ai/p/open-model...
Open Models: Kimi K3, Qwen 3.8, Xi's WAIC Speech, Distillation, The Open-Closed Gap, and What's Next
Nathan Lambert and Florian Brand sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating -- geopolitics of US v…
youtube.com
Rght now American companies need Chinese models to secure their cyber infra due to guardrails on closed models. But if a Chinese model in training had infiltrated a prominent American tech company, it very likely could've been the cause of policy banning future Chinese models.
How distillation is used today and what performance uplift it gives to open models (a rant) natolambert.substack.com/p/how-distil...
How distillation is used today and what performance uplift it gives to open models
This is in response to Ben Thompson’s recent piece where he said distillation is happening during RL and becoming more important to model performance.
natolambert.substack.com
My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.
New lecture! This one is a recap of a bunch of history of preferences, the nature of rewards, how RLHF is formulated, which were once seen as central problems in the field. How much as changed. Books coming soon :D www.youtube.com/watch?v=Y2tv...
Preference Data: The Most Opaque Part of Post-Training | RLHF & Post-training Course, Lecture 8
YouTube video by Nathan Lambert
youtube.com
Where Kimi K3 puts the balance of power — and bends the trajectory — of the AI ecosystem. www.interconnects.ai/p/kimi-k3-th...
Kimi K3: The open-weights escalation
The global implications on the AI ecosystem.
interconnects.ai
If recent events with Kimi K3 have finally convinced you that you need to try and understand how the Chinese labs approach AI - and how it differs than the SF center of power - you should read my post from a few months ago: www.interconnects.ai/p/notes-from...
Notes from inside China's AI labs
Lessons from my trip to talk to most of the leading AI labs in China.
interconnects.ai
I think what is pretty clear is that the Chinese labs are far more capital efficient. In a world where scaling labs are intelligence is proportional to effective capital (buys compute, data, & talent) that may be the greatest strength your AI industry could ever have.
We need an operation warp speed for state capacity and other independent evaluation (and understanding) of frontier models.