Announcing Olmo 3, a leading fully open LM suite built for reasoning, chat, & tool use, and an open model flow—not just the final weights, but the entire training journey. Best fully open 32B reasoning model & best 32B base model. 🧵
Noah A. Smith
@nlpnoah.bsky.social
Researcher in NLP, ML, computer music. Prof @uwcse @uwnlp & helper @allen_ai @ai2_allennlp & familiar to two cats. Single reeds, tango, swim, run, cocktails, מאַמע־לשון, GenX. Opinions not your business.
"We are committed to our fully open ethos. That's why we release everything—weights, code, training data, checkpoints, all of it." — @nlpnoah.bsky.social at the Madrona IA Summit last week.
#UWAllen prof @nlpnoah.bsky.social will be the new vice provost for artificial intelligence and Charles & Lisa Simonyi Endowed Chair for Artificial Intelligence & Emerging Technologies thanks to a $10M gift supporting @uofwa.bsky.social's global leadership in #AI. www.washington.edu/news/2025/11...
$10M gift from Charles and Lisa Simonyi establishes AI@UW to advance artificial intelligence and emerging technologies
The University of Washington today announced a foundational $10 million gift from philanthropists Charles and Lisa Simonyi to support groundbreaking work in artificial intelligence and emerging...
washington.edu
a suicide note: I only logged in 200 times because your shitty system kept logging me out now a simple reviewing task has become endless I cannot do this again farewell @openreview.bsky.social
RewardBench 2 is here! We took a long time to learn from our first reward model evaluation tool to make one that is substantially harder and more correlated with both downstream RLHF and inference-time scaling.
I am furious about the news out of Colorado of Jews being lit on fire for the crime of calling for hostages to be released. I am furious at how dirtbag leftists on this website and elsewhere treat this as a game where violence against American Jews is justified if we aren't good tokens.
Congratulations to @yizhongw.bsky.social, who successfully defended his PhD thesis today! Advised by @hanna-nlp.bsky.social and me.
trying to decide whether to go to Vienna in late July (there's a formerly-AI conference happening there and I might want to see some friends who still go) maybe you can help me decide who's going to be there and wants to rent a place and play music together
(I don't know why this will never get old) *upload photo of anyone at all to Molmo, type "where is the scientist?"* Molmo: I don't see a scientist here
the economist called us losers and we ... rolled our eyes because their argument is behind a paywall Gen X (at its best) doesn't feed trolls
Why Gen X is the real loser generation
Don’t cry for millennials or Gen Z. Save your pity for those in their 50s
economist.com
used a LM to diagnose a medical problem, figure out steps to get treatment with minimal phone calls and time wasted waiting for healthcare providers to perform administrative tasks (1/n)
intellectually I get that his name is "Robert Prevost" but in my head I can't not hear it as "Ford Prefect"
before we get to their math, can someone explain how modern politicians choose to capitalize words? is there any rhyme or reason to it, or is it just a new version of boomer-random?
By this Pam calculation, one pill has to kill more than five people. Math is hard.
"When we erode the boundaries between the academic and the political, we ultimately harm both."
Opinion | How About We Don’t Bring Our Whole Selves to Work?
Politics has no place at universities or in the classroom.
nytimes.com
always hold these two thoughts in your mind when considering evaluation of AI systems: 1. rigorous evaluation is essential; we must not go back to the 1980s (anecdote world) 2. continuous improvement to eval methodology is essential; evaluations are usually broken in ways worth fixing
1/ Science is only as strong as the benchmarks it relies on. So how fair—and scientifically rigorous—is today’s most widely used evaluation benchmark? We took a deep dive into Chatbot Arena to find out. 🧵
protip: the less you reply, the less email you get
Yes writing emails is a pain so instead of hiding behind an AI to do it for you like a coward you need to simply have the courage of your convictions and stand up for what matters and simply not reply to them like the rest of us
True story: the NLP community elected the person holding this position to the presidency of our professional society. Let's take the greatest, most impactful thing to come out of our little world and shit all over it! It's the Tr*mp presidency in miniature
🔭 Science relies on shared artifacts collected for the common good. 🛰 So we asked: what's missing in open language modeling? 🪐 DataDecide 🌌 charts the cosmos of pretraining—across scales and corpora—at a resolution beyond any public suite of models that has come before.
Ever wonder how LLM developers choose their pretraining data? It’s not guesswork— all AI labs create small-scale models as experiments, but the models and their data are rarely shared. DataDecide opens up the process: 1,050 models, 30k checkpoints, 25 datasets & 10 benchmarks 🧵
Meet Ai2 Paper Finder, an LLM-powered literature search system. Searching for relevant work is a multi-step process that requires iteration. Paper Finder mimics this workflow — and helps researchers find more papers than ever 🔍
This is just to say I have covered your Tesla in Kraft Singles which probably irritated and perplexed you Do not forgive me I will do it again and again and again
a small change to building your BPE tokenizer gets your pretrained LM 8 MMLU points (for example) and 27% inference-time efficiency boost ...
We created SuperBPE🚀, a *superword* tokenizer that includes tokens spanning multiple words. When pretraining at 8B scale, SuperBPE models consistently outperform the BPE baseline on 30 downstream tasks (+8% MMLU), while also being 27% more efficient at inference time.🧵
It’s unfair to compare the trashing of Teslas to the Boston Tea Party because tea is worth something
Humbled to make @fastcompany.com's 2025 most innovative companies list for making AI models that are truly open. "Ai2 is setting a strong benchmark for what transparency can look like in the entire AI industry."🎉
Senate Democrats who voted YES on cloture for the CR: Schumer Gillibrand Fetterman Schatz Durbin King Shaheen Hassan Peters Cortez Masto