🚋 New blog post: On "infinite" learning-rate schedules and how to construct them from one checkpoint to the next. fabian-sp.github.io/posts/2025/0...
Infinite Schedules and the Benefits of Lookahead
TL;DR: Knowing the next training checkpoint in advance (“lookahead”) helps to set the learning rate. In the limit, the classical square-root schedule appears on the horizon.
fabian-sp.github.io