Kyle Lo @ ICML2026 🇰🇷

@kylelo.bsky.social

language models, data & evals, prev co-lead of Olmo @ai2.bsky.social, nlp @uwcse, statistics @uw, open science, tabletop, seattle, he/him,🧋 kyleclo.com

New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!

Figure 1: A web passage scored across our 12 narrative dimensions. Agency and setting dimensions are rated on a 1–5 Likert scale, temporal sequencing and causal density are passage-level proportions (0–1), and event density is the rate of event triggers per token. This passage scores high on agency and event features but low on setting, a “narrative profile” commonly seen across first-person web narratives.Figure 6: UMAP reduction of SBERT embeddings for 20,000 randomly sampled NARRADOLMA documents, colored by PC1 score (interiority). Labels are based on manual examination. Overlays for all three PCs appear in Fig. A5.
Teagan Johnson@teagrjohnson.bsky.social · 2mo ago

1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:

This image depicts the proportion of each Dolma category in the top quartile for the first three principal components.

during in Olmo 3 we thought long context is just finding good data nope! model architecture matters & it's hard to recover if mess it up led by @abertsch.bsky.social, we release many pretrain runs w/ small arch changes and show huge long context performance diffs

Bild
Ai2@ai2.bsky.social · 3mo ago

Recipes for teaching language models to handle long inputs don't work equally well across model families. We wanted to know why—is it the architecture, the training data, or both? 🧵

our new Olmo Hybrid model combines attention with linear RNN layers 🍣training efficiency is crazy good. the model reaches same MMLU score as Olmo 3 in 50% of the tokens. also see this in many other tasks as always: weights, data, ckpts, training code, etc. all fully open

Bild
Ai2@ai2.bsky.social · 5mo ago

Introducing Olmo Hybrid, a 7B fully open model combining transformer and linear RNN layers. It decisively outperforms Olmo 3 7B across evals, w/ new theory & scaling experiments explaining why. 🧵

DrawEduMath is our benchmark testing VLM understanding of K-12 student math work, which is prerequisite for their use in educational contexts one year after, while VLMs are strong math solvers today, they still underperform on our bench, esp for students who need the most help

Lucy Li@lucy3.bsky.social · 5mo ago

Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵

Title, author list, and two figures from the paper. 
Title: The Aftermath of DrawEduMath: Vision Language Models
Underperform with Struggling Students and Misdiagnose Errors
Authors: Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo
Figure 1: On the left is a math problem, where students are asked to draw x < 5/2 on a number line. The right side shows two example student responses that differ in correctness. DrawEduMath pairs each math problem with one student response, and prompts VLMs to answer questions about the student response.
Figure 2: VLMs consistently perform worse on answering DrawEduMath benchmark questions pertaining to erroneous student responses. Performance on non-erroneous student responses is labeled with specific VLMs’ names; that same model’s performance on erroneous student responses is directly below.

our paper on data mixing for LMs is out! while building Olmo 3, we saw gaps between data mixing literature and real practice 🐠choosing proxy size, # runs, sampling, regression, constraints.. 🐟data shifts during LM dev: can we reuse past experiments? Olmix tackles them all!

Ai2@ai2.bsky.social · 6mo ago

Data mixing – determining how much web text, code, math, etc., you need for LM development – is a first-order lever on model quality. Introducing Olmix: a framework for configuring mixing methods at the start of dev & efficiently updating as data changes throughout. 🧵

some thoughts about skill degradation w/ AI coding im onboard w views that "english is the new programming language" & "software engineering", translating ambiguous goals to technical specs/execution, is still a skill. im more concerned w shift from my role as a writer to a reviewer and

using opus to extract research topics from papers & it was giving me useless words like "training", "datasets", and "evaluation" kept prompting it w examples of more informative topics and it ended up with "LLM training", "LLM datasets", and "LLM evaluation" thx

bsky wish list i like the idea of different feeds but i actually want my subscription to select feeds to be taken as a preference signal ("more like this") that informs a "home/default" feed. i really dislike the UX of having to tab through each subscribed feed, esp when there's also post overlap

during neurips, we kept the RL run going & model kept getting better 😂 Olmo 3.1 is a.. 🐡 32B Thinking, still best fully-open model to-date 🐠 32B Instruct, for ppl who hate long yapping, as good as qwen3 we added 10 more pages to the paper! thx for community feedback from convos at neurips

Ai2@ai2.bsky.social · 8mo ago

Olmo 3.1 is here. We extended our strongest RL run and scaled our instruct recipe to 32B—releasing Olmo 3.1 Think 32B & Olmo 3.1 Instruct 32B, our most capable models yet. 🧵

we released Olmo 3! lot of exciting stuff but wanna focus on: 🐟Olmo 3 32B Base, the best fully-open base model to-date, near Qwen 2.5 & Gemma 3 on diverse evals 🐠Olmo 3 32B Think, first fully-open reasoning model approaching Qwen 3 levels 🐡12 training datasets corresp to different staged training

BildBild

not happy abt gpt 5.1 update. it's making way more mistakes compared to gpt 5 on basic stuff latex table formatting errors (straight up missing "&" so columns misaligned, or dropping a whole column, or shifting values by 1 position), feels unusable imo 😒

why intern at Ai2? 🐟interns own major parts of our model development, sometimes even leading whole projects 🐡we're committed to open science & actively help our interns publish their work reach out if u wanna build open language models together 🤝 links 👇

Bild