Nathan

@saylortwift.hf.co

ML engineer at @huggingface 🤗, Evaluation, Open LLM Leaderboard and lighteval

Evaluation was just made easier 💯 We merged a huge refacto of lighteval making easier to add: 🔄 Multiturn tasks 🖼️ Multimodal tasks 📝 Plus unified logs for thorough benchmark analysis Benchmarks guys, what evals would you like to see added ?

🔥 Evaluating LLMs? You need Lighteval — the fastest, most flexible toolkit for benchmarking models, built by @huggingface Now with: ✅ Plug & play custom model inference (evaluate any backend) 📈 Tasks like AIME, GPQA:diamond, SimpleQA, and hundreds more Details below 🧵👇

openai really has some nice benchmarks, one of them being simpleqa. a simple fact-checking benchmark, short questions and straight answers i've been using @huggingface's lighteval and inference providers and litellm to evaluate all those models in less than a few hours 🤩 1/N

Bild

🚀 Introducing ✨ YourBench ✨ ! Build custom evals instantly using your private docs & see how your custom fine-tuned models perform on your unique tasks. Congrats to @sumukx @clefourrier and @ailozovskaya for their incredible work ! Game-changing for LLM evaluation 🚀 1/2

Bild

Just wrapped up evaluations on @deepseek_ai's V3 0324! 🚀 Impressive gains in math and GPQA, but instruction following took a slight hit. More concerning—AIME25 remains unchanged. Possible contamination issues? 🤔

Bild

WOW. The Qwen team did NOT come to play.🔥 Just look at these insane results from the OpenEval team—absolutely impressive. Huge congrats! 👏 @Alibaba_Qwen

Bild

Everyone's talking about GPT-4.5 quality, so we ran benchmarks! Did NOT expect it to be such a leap from GPT-4o—now on par with Claude 3.7 and even ahead of DeepSeek Llama 70B (a thinking model!). Congrats to the team @OpenAI !

Bild

Everyone's talking about GPT-4.5 quality, so we ran benchmarks! Did NOT expect it to be such a leap from GPT-4o—now on par with Claude 3.7 and even ahead of DeepSeek Llama 70B (a thinking model!). Congrats to the team @OpenAI ! Now open-source it and drop it on the Hub 🤗

Bild

we just reproduced Claude 3.7 results for you 📈 TLDR: we get what they announced. We also used AIME 2025 to test for contamination on the 2024 version and score are similar on both benchmarks ! Great job to the @AnthropicAI team ! More details in thread 👇 1/3

Bild

Today marks my 2 years at @huggingface! Time flies !! Working with those people for 2 years now, I can tell you there is no better place to build ethical, open AI. Hf folks are both kind and incredibly talented, I can't wait to work on many more exciting projects with them 🤩

DeepSeek R1 continues to impress! I just integrated the Olympiad Bench— a collection of elite-level Chinese and English scientific problems— into LightEval and tested GPT-4o against R1. The results are insane. Full details + how to reproduce in the thread 👇

Bild

Making SmolLM2 more reproducible: open-sourcing our training & evaluation toolkit 🛠️ github.com/huggingface/... Pre-training & evaluation code, synthetic data generation pipelines, post-training scripts, on-device tools & demos Apache 2.0. V2 data mix coming soon! Which tools should we add next?

GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models

Everything about the SmolLM & SmolLM2 family of models - GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models

github.com

Check out how easy it is to do LLM evals with LightEval! * any dataset on the 🤗 Hub can become an eval task in a few lines of code: customize the prompt, metrics, parsing, few-shots, everything! * model- and data-parallel inference * auto batching with the new vLLM backend

A screenshot of LightEval benchmarking results in a terminal

It's "on-device LLM" today. Soon, it'll be "on-chip" LLM. Or LLM cores. The system default local LLM. The coding framework's default local LLM. I find this incredibly exciting. A privacy-first, self-contained, user-owned AI—a 24/7 agent for action, insights & feedback. github.com/huggingface/...

GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models

Everything about the SmolLM & SmolLM2 family of models - GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models

github.com

This week (ish) in 🌤️ LLM evaluation 🔥 📊 A statistical approach to model evaluation @AnthropicAI 📐 Frontier MATH: a benchmark for evaluating advanced Mathematical reasoning in AI @EpochAIResearch 📝 Say What You Mean: A Response to 'Let Me Speak Freely' @dottxtai 🧵 👇

Bild