Nathan Lambert

@natolambert.bsky.social

A LLN - large language Nathan - (RL, RLHF, society, robotics), athlete, yogi, chef Writes http://interconnects.ai Prev Ai2/Olmo, HuggingFace, Berkeley, and normal places

If you're teaching a class on post-training (or part of a course) and my book, slides, code or videos don't help you, please lmk how I can improve it! Has been a ton of work getting everything done and part of the ROI is hope that it helps more education work grows around LLMs.

The pace of progress on models from so many organizations at once is genuinely incredible. Building LLMs isn't driven by rare secrets, but consistent effort, mass capital, and effective organization design. It is great that know-how of such a powerful technology is diffused.

Another new lecture! Lecture 11 is a tool-use/function calling/agentic 101. I almost skipped this one, as this chapter started as the only skill-specific topic in the book, but since writing it tool-use has only become more foundational to modern models.

Bild

Frontier labs will be viable businesses by being able to integrate and optimize inference at lower cost/performance than most other models. They'll have a margin advantage on open models for the foreseeable future. This is one of those flex's imo and I doubt they're losing money.

Bild

Lecture 10 of my course! Nominally on regularization in RL, so I discuss the evolving role of the KL penalty in RL, but also a set of nice RL papers that explain what RL helps models generalize better than SFT -- with theory supporting it.

Bild

Kimi K3 with more likes than downloads on HuggingFace is definitely showing us a glimpse of the future on open models. It's way less about individual access, and more of a distributed platform layer for companies.

Bild

Making talks with AI agents is awesome. I just told Fable to make a slide with real data on the KL distance from one of our reference Olmo 2 models and it made this with the wandb api (I edited text slightly).

Bild

New (shorter) lecture! Over-optimization, foundations of reward hacking, sycophancy, verbosity, etc. In recording this, I realized that rubrics are going to be prone to overopt in a way like reward models, where RLVR is its own thing. Fundamentals, history, and reflections! youtu.be/y04JhXpiI4s

Over-Optimization and RLHF’s Bad Reputation | Post-Training Course, Lecture 9

In this lecture, we discuss how RLHF got a bad reputation, how over-optimization compares to over-fitting, and the nuance in getting preference tuning right. It's a shorter one, in order to keep…

youtube.com

At night I dream of a distillation debate grounded in public, technical info, not reading tea leaves of backroom deals and political intrigue. Then I wake up and I'm crushed by reality of chaos, potentially classified information, and a spiraling global AI ecosystem.

Rght now American companies need Chinese models to secure their cyber infra due to guardrails on closed models. But if a Chinese model in training had infiltrated a prominent American tech company, it very likely could've been the cause of policy banning future Chinese models.

My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.

Bild

I think what is pretty clear is that the Chinese labs are far more capital efficient. In a world where scaling labs are intelligence is proportional to effective capital (buys compute, data, & talent) that may be the greatest strength your AI industry could ever have.