Olmo 3.1: even more RL = even more RL-Zero! @saurabhshah2.bsky.social and I tweaked some hyperparams and prompts, @hamishivi.bsky.social and @finbarr.bsky.social improved the code and boom! New Olmo 3.1 RL-Zero 👾 An updated, solid baseline for your RL and reasoning research
Olmo 3.1 32B Think shows that not just frontier labs can scale RL. My favorite RL run yet over 7+ years of doing RL. The biggest fully open RL run ever? Gold stars on downstream evals is our original release, this latest one is the final checkpoint on the plot.