Quanquan Gu

@quanquangu.bsky.social

Professor @UCLA, Research Scientist @ByteDance | Recent work: SPIN, SPPO, DPLM 1/2, GPM, MARS | Opinions are my own

This Thanksgiving, I want to express my heartfelt gratitude to all the students, colleagues, and collaborators who have contributed to the success of SPIN, SPPO, DPLM, GPM, MARS, and many other projects. Your hard work and dedication continue to be truly inspiring.

Graduate school application season is here again, and I just spent the whole weekend writing a few letters! 🎓 When submitting, many schools still require overly detailed forms filled with endless questions and checkboxes. Isn’t it time to simplify this process and focus on what truly matters?

It's Sunday morning so taking a minute for a nerdy thread (on math, tokenizers and LLMs) of the work of our intern Garreth By adding a few lines of code to the base Llama 3 tokenizer, he got a free boost in arithmetic performance 😮 [thread]

Bild

I’m increasingly unsure there are specific rules/laws for pretraining schedule except that peak LR instability means that you need a warmup and a cooldown. And in the post-Chinchilla era you probably want to slow cook the model (so go easy on lr, smooth the schedule).

Everything you want to know about FSDP and more... FSDP works on layers at a time and enables perf wins by overlapping the data-fetching part of the next layer with the computation of the current layer. This overlap means that the data-fetching can be “hidden”.

Bild