Noah A. Smith

@nlpnoah.bsky.social

Researcher in NLP, ML, computer music. Prof @uwcse @uwnlp & helper @allen_ai @ai2_allennlp & familiar to two cats. Single reeds, tango, swim, run, cocktails, מאַמע־לשון, GenX. Opinions not your business.

Announcing Olmo 3, a leading fully open LM suite built for reasoning, chat, & tool use, and an open model flow—not just the final weights, but the entire training journey. Best fully open 32B reasoning model & best 32B base model. 🧵

Bild

RewardBench 2 is here! We took a long time to learn from our first reward model evaluation tool to make one that is substantially harder and more correlated with both downstream RLHF and inference-time scaling.

The RewardBench 2 Leaderboard on HuggingFace.

I am furious about the news out of Colorado of Jews being lit on fire for the crime of calling for hostages to be released. I am furious at how dirtbag leftists on this website and elsewhere treat this as a game where violence against American Jews is justified if we aren't good tokens.

trying to decide whether to go to Vienna in late July (there's a formerly-AI conference happening there and I might want to see some friends who still go) maybe you can help me decide who's going to be there and wants to rent a place and play music together

used a LM to diagnose a medical problem, figure out steps to get treatment with minimal phone calls and time wasted waiting for healthcare providers to perform administrative tasks (1/n)

always hold these two thoughts in your mind when considering evaluation of AI systems: 1. rigorous evaluation is essential; we must not go back to the 1980s (anecdote world) 2. continuous improvement to eval methodology is essential; evaluations are usually broken in ways worth fixing

Marzieh Fadaee@mziizm.bsky.social · last yr.

1/ Science is only as strong as the benchmarks it relies on. So how fair—and scientifically rigorous—is today’s most widely used evaluation benchmark? We took a deep dive into Chatbot Arena to find out. 🧵

True story: the NLP community elected the person holding this position to the presidency of our professional society. Let's take the greatest, most impactful thing to come out of our little world and shit all over it! It's the Tr*mp presidency in miniature

Bild

🔭 Science relies on shared artifacts collected for the common good. 🛰 So we asked: what's missing in open language modeling? 🪐 DataDecide 🌌 charts the cosmos of pretraining—across scales and corpora—at a resolution beyond any public suite of models that has come before.

Ai2@ai2.bsky.social · last yr.

Ever wonder how LLM developers choose their pretraining data? It’s not guesswork— all AI labs create small-scale models as experiments, but the models and their data are rarely shared. DataDecide opens up the process: 1,050 models, 30k checkpoints, 25 datasets & 10 benchmarks 🧵

Plot shows the relationship between compute used to predict a ranking of datasets and how accurately that ranking reflects performance at the target (1B) scale of models pretrained from scratch on those datasets.

Meet Ai2 Paper Finder, an LLM-powered literature search system. Searching for relevant work is a multi-step process that requires iteration. Paper Finder mimics this workflow — and helps researchers find more papers than ever 🔍

Screenshot of the Ai2 Paper Finder interface

This is just to say I have covered your Tesla in Kraft Singles which probably irritated and perplexed you Do not forgive me I will do it again and again and again

Screenshot of a Reddit post in r/Seattle

To My Neighbor Whose Tesla Is Covered in Kraft Singles
1. I am the one who keeps doing this
2. This is not because you own a Tesla, but because of who you are as a person and the choices you have made.
3. Every time you veer out of your way to splash people while we are waiting for the bus, I will do it again.
4. You are never going to catch me.

That is all.

Humbled to make @fastcompany.com's 2025 most innovative companies list for making AI models that are truly open. "Ai2 is setting a strong benchmark for what transparency can look like in the entire AI industry."🎉

Senate Democrats who voted YES on cloture for the CR: Schumer Gillibrand Fetterman Schatz Durbin King Shaheen Hassan Peters Cortez Masto