Quentin Gallouédec

@qgallouedec.hf.co

PhD - Research @hf.co 🤗 TRL maintainer

🚀 TRL 0.14 – Featuring GRPO! 🚀 TRL 0.14 brings *GRPO*, the RL algorithm behind 🐳 DeekSeek-R1 . ⚡ Blazing fast generation with vLLM integration. 📉 Optimized training with DeepSpeed ZeRO 1/2/3.

Bild

[Stonks] TRL is a Python library for training language models. It has seen impressive growth this year. Lots of new features, an improved codebase, and this has translated into increased usage. You can count on us to do even more in 2025.

Bild

🚨 TRL 0.13 is out! 🤗 Featuring a Process-supervised Reward Models (PRM) Trainer 🏋️ PRMs empower LLMs to "think before answering"—a key feature behind OpenAI's o1 launch just two weeks ago. 🚀

Bild

We outperform Llama 70B with Llama 3B on hard math by scaling test-time compute 🔥 How? By combining step-wise reward models with tree search algorithms :) We're open sourcing the full recipe and sharing a detailed blog post 👇

Bild

Join us at Hugging Face as an intern if you want to contribute to amazing open-source projects, and develop LLM's best finetuning library, aka TRL. 🧑‍💻 Full remote 🤯 Exciting subjects 🌍 Anywhere in the world 🤸🏻 Flexible working hours Link to apply in comment 👇

Bild

It's Sunday morning so taking a minute for a nerdy thread (on math, tokenizers and LLMs) of the work of our intern Garreth By adding a few lines of code to the base Llama 3 tokenizer, he got a free boost in arithmetic performance 😮 [thread]

Bild