Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization. Led by Sara Dragutinović and advised by Rajesh Ranganath arxiv.org/abs/2603.00742
Yedi Zhang
@yedizhang.bsky.social
PhD student @ Gatsby Unit UCL http://yedizhang.github.io/
Come chat about this @iclr-conf.bsky.social! Friday 3:15 PM, Pavilion 4, Poster #4216
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures? Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham arxiv.org/abs/2512.20607
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures? Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham arxiv.org/abs/2512.20607
How does in-context learning emerge in attention models during gradient descent training? Sharing our new Spotlight paper @icmlconf.bsky.social: Training Dynamics of In-Context Learning in Linear Attention arxiv.org/abs/2501.16265 Led by Yedi Zhang with @aaditya6284.bsky.social and Peter Latham