With other folks at 🍏, @brunokm.bsky.social has worked on a complete(d) parameterisation for NNs that can *transfer* locally tuned hyperparameters: tune optimizers' parameters (e.g. LR) *per module/depth* using an evolutionary search on small models → they transfer perf. gains to much larger models
In our new work — Complete(d)P — we try to answer 3 questions about hyperparameter (HP) scaling: ● How to transfer across model size, tokens&batch-size?→ Complete(d)P ● Do per-module HPs matter? ✔️2x speed-ups possible ● Do they transfer to larger scale? ✔️ With the right parameterisation