Pierre Chambon

@pierrechambon.bsky.social

PhD at FAIR (Meta) and INRIA Former researcher at Stanford University

Llama 4 results out on ✨BigO(Bench)✨! Llama 4 Maverick is top 4 all@1 on Time Complexity Generation and top 2🥈coeffFull on Time Complexity Ranking (beating R1, though not using any reasoning tokens). The model is less performant on Space Complexity. 👇All links below👇

Bild

✨BigO(Bench)✨ Leaderboard Update! 3 models added to our benchmark: 🏆 nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 🧑‍💻 agentica-org/DeepCoder-14B-Preview 🤲 all-hands/openhands-lm-32b-v0.1 Thanks @vllm_project and @huggingface for quickly supporting inference! 👇All links below👇

Bild

Does your LLM truly comprehend the complexity of the code it generates? 🥰 Introducing our new non-saturated (for at least the coming week? 😉) benchmark: ✨BigO(Bench)✨ - Can LLMs Generate Code with Controlled Time and Space Complexity? Check out the details below !👇

Bild