excited to see frens at #icml2026 & present 🐟 Olmix: efficient data mixing under token constraints & evolving data domains 🐡 How2Everything: mining the web for diverse procedural tasks for train & eval 🐠 happy to chat data & evals, both pre & post-training
Kyle Lo @ ICML2026 🇰🇷
@kylelo.bsky.social
language models, data & evals, prev co-lead of Olmo @ai2.bsky.social, nlp @uwcse, statistics @uw, open science, tabletop, seattle, he/him,🧋 kyleclo.com
audiobooks have rlly improved my new commute. discovered Libby & dunno why anyone would pay for Audible
Excited to share this. @neel2112.bsky.social, @mariaa.bsky.social, and I analyzed 500K anonymous ChatGPT convos (shared w/ consent from WildChat) to see if people were generating fiction. We found tons of stories, fanfiction & erotica. Many users iterated on the same stories for days and weeks.
I'm really sorry to miss all the fun at @facct.bsky.social this year! But @nlp-amelie.bsky.social and Mattes Ruckdeschel, the first two authors of this work ⬇️, are around.
A Human-Centric Framework for Data Attribution in LLMs (FaCCT'26) TLDR: LLMs screwed up data economy, and NLP should help to re-design incentives with data attribution. Here's a moonshot for an alternative LLM use paradigm, for the users, creators & intermediaries. arxiv.org/abs/2602.10995 /1
New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!
1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:
during in Olmo 3 we thought long context is just finding good data nope! model architecture matters & it's hard to recover if mess it up led by @abertsch.bsky.social, we release many pretrain runs w/ small arch changes and show huge long context performance diffs
Recipes for teaching language models to handle long inputs don't work equally well across model families. We wanted to know why—is it the architecture, the training data, or both? 🧵
Today I'm saying farewell to @ai2.bsky.social. I'm so proud of our team & grateful to have shared fully-open Olmo, Dolma, olmOCR, Molmo, etc with the world I know the team is more committed than ever to advancing open-source & open-science. Forever rooting for my dear friends 🫶
our new Olmo Hybrid model combines attention with linear RNN layers 🍣training efficiency is crazy good. the model reaches same MMLU score as Olmo 3 in 50% of the tokens. also see this in many other tasks as always: weights, data, ckpts, training code, etc. all fully open
Introducing Olmo Hybrid, a 7B fully open model combining transformer and linear RNN layers. It decisively outperforms Olmo 3 7B across evals, w/ new theory & scaling experiments explaining why. 🧵
DrawEduMath is our benchmark testing VLM understanding of K-12 student math work, which is prerequisite for their use in educational contexts one year after, while VLMs are strong math solvers today, they still underperform on our bench, esp for students who need the most help
Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵
our paper on data mixing for LMs is out! while building Olmo 3, we saw gaps between data mixing literature and real practice 🐠choosing proxy size, # runs, sampling, regression, constraints.. 🐟data shifts during LM dev: can we reuse past experiments? Olmix tackles them all!
Data mixing – determining how much web text, code, math, etc., you need for LM development – is a first-order lever on model quality. Introducing Olmix: a framework for configuring mixing methods at the start of dev & efficiently updating as data changes throughout. 🧵
incredibly fun project led by our intern yapei chang we mined the web for thousands of real-world “how to do X” step by step instructions and turned it into a dataset, synth data training procedure, eval suite, etc.
LLMs often generate step-by-step instructions, from real-world tasks (how do I file taxes?) to plans for AI agents. Improving this is hard: outputs can sound fluent for steps that don't work, and current datasets cover few domains. How2Everything evals/trains for this at scale. 🧵
our open model proving out specialized rag LMs over scientific literature has been published in nature ✌🏻 congrats to our lead @akariasai.bsky.social & team of students and Ai2 researchers/engineers www.nature.com/articles/s41...
0 days since last mixup of eval results between "copa" (choice of plausible alternatives) & "coqa" (conversational QA) tasks 😐
The 5th Generation, Evaluation, and Metrics (GEM) Workshop will be at #ACL2026! Call for papers is out. Topics include: 🐟 LMs as evaluators 🐠 Living benchmarks 🍣 Eval with humans and more New for 2026: Opinion & Statement Papers! Full CFP: gem-workshop.com/call-for-pap...
some thoughts about skill degradation w/ AI coding im onboard w views that "english is the new programming language" & "software engineering", translating ambiguous goals to technical specs/execution, is still a skill. im more concerned w shift from my role as a writer to a reviewer and
lucky to chat w sen. patty murray about olmo & importance of fully open AI
using opus to extract research topics from papers & it was giving me useless words like "training", "datasets", and "evaluation" kept prompting it w examples of more informative topics and it ended up with "LLM training", "LLM datasets", and "LLM evaluation" thx
just realized ive had food on my face all day & nobody at office told me, thx ai2 frens 😫
bsky wish list i like the idea of different feeds but i actually want my subscription to select feeds to be taken as a preference signal ("more like this") that informs a "home/default" feed. i really dislike the UX of having to tab through each subscribed feed, esp when there's also post overlap
just had hechalou’s yin yang milk tea and i think i’ve transcended 🤤
during neurips, we kept the RL run going & model kept getting better 😂 Olmo 3.1 is a.. 🐡 32B Thinking, still best fully-open model to-date 🐠 32B Instruct, for ppl who hate long yapping, as good as qwen3 we added 10 more pages to the paper! thx for community feedback from convos at neurips
Olmo 3.1 is here. We extended our strongest RL run and scaled our instruct recipe to 32B—releasing Olmo 3.1 Think 32B & Olmo 3.1 Instruct 32B, our most capable models yet. 🧵
I'll be at #NeurIPS2025 from Tues-Sat! Come say hi 👋 if you wanna chat about 🦈 olmo 3 stories 🐟 pretraining data & evals 🍣 midtraining shouldnt exist 🐠 model specialization 🐡 AI for education 🍥 tabletop games
we released Olmo 3! lot of exciting stuff but wanna focus on: 🐟Olmo 3 32B Base, the best fully-open base model to-date, near Qwen 2.5 & Gemma 3 on diverse evals 🐠Olmo 3 32B Think, first fully-open reasoning model approaching Qwen 3 levels 🐡12 training datasets corresp to different staged training
not happy abt gpt 5.1 update. it's making way more mistakes compared to gpt 5 on basic stuff latex table formatting errors (straight up missing "&" so columns misaligned, or dropping a whole column, or shifting values by 1 position), feels unusable imo 😒
picking between 3 checkpoints w/ same benchmark scores but what if one of them is agi
why intern at Ai2? 🐟interns own major parts of our model development, sometimes even leading whole projects 🐡we're committed to open science & actively help our interns publish their work reach out if u wanna build open language models together 🤝 links 👇
congrats to our olmo earth team 🌎 small multimodal foundation language models + system for finetuning for important uses like agriculture, wildfire management, conservation & more 🌿
Introducing OlmoEarth 🌍, state-of-the-art AI foundation models paired with ready-to-use open infrastructure to turn Earth data into clear, up-to-date insights within hours—not years.