@kylelwiggers.bsky.social

Ai2 Comms Lead | kylew@allenai.org | Pronouns: he/him

We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers. It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes. Here's what they told us. 🧵

Our @skylightmarine.bsky.social team is rolling out Shippy, an AI agent for the people protecting our ocean in real time. A wrong answer at these stakes can send a patrol vessel miles off course. Here's the architecture we built so analysts can trust it. 🧵

Bild

olmOCR 2 is now in the Ai2 Playground—our home for our fully open text, video, and image understanding models. 🧵

Bild

The Danish Foundation Models (DFM) project is adapting our modular FlexOlmo architecture into a lighter-weight system that runs on commodity hardware—putting collaborative model building within reach of smaller research groups & organizations. 🧵

Bild

Today we're releasing OlmoEarth v1.2, the latest in our family of open foundation models for Earth observation. 🌍 We've switched to rotary positional embeddings (RoPE), which reduces artifacts in the embeddings & gives a small performance boost. 🧵

Bild

Hybrid (transformer–RNN) models are fast becoming a serious alternative to the transformer, but a big question remains: how do they process tokens differently & how does this impact performance? We compared our transformer (Olmo 3) & hybrid (Olmo Hybrid) models to find out. 🧵

Bild

We're releasing MolmoMotion, a 3D motion forecasting model. Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. 🧵

Building an LLM means evaluating it over & over as it changes. Tweak a hyperparameter or scale the model up, & every new checkpoint sends you back through the same benchmarking loop. We're releasing olmo-eval, a workbench built for this kind of iterative model development. 🧵

Bild

LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work. So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560 We made ModSleuth to trace this. 🧵

Bild

Today we’re releasing OlmoEarth v1.1. It’s 3x cheaper to run than v1 while delivering the same state-of-the-art performance—and fully open. 🧵

Bild

Brendan Works is a product manager focused on paratransit services in Seattle. See how he built PointCheck, a website accessibility checker powered by our open Molmo, MolmoWeb, & Olmo 3 models. 👇

Bild

Now available in AstaLabs in limited research preview: MyScholarQA, a personalized version of ScholarQA for scientific deep research. ScholarQA helps synthesize evidence from 12M+ open-access papers. MyScholarQA adds user profiles to tailor that synthesis to you. 🧵

Artificial Analysis relies on our IFBench eval to test how closely models follow user prompts. Most evals in their Intelligence Index saturate within months. IFBench hasn't because it measures what others miss—and what frontier models still struggle with. 🧵

Bild

Today we’re releasing EMO, a new mixture-of-experts (MoE) model trained so modular structure emerges directly from data without human-defined priors. EMO can use a small subset of its experts for a given task while keeping near full-model performance. 🧵

Bild

Robotics models often struggle outside controlled environments. Ours is built to work in real ones. Today we're launching MolmoAct 2, which can assist with a host of chores & lab tasks, plus the MolmoAct 2-Bimanual YAM dataset—the largest open robotics dataset of its kind. 🧵

Today we published a Q&A with Interim CEO Peter Clark on what’s next for Ai2, from advancing truly open AI systems to applying AI in areas like scientific discovery & the planet. 👇

Bild

New AstaBench results show frontier models making progress on scientific research, but the benchmark remains far from solved. Claude Opus 4.7 leads overall at 58.0%, while GPT-5.5 comes within 5.1 points at less than half the measured cost per problem. 🧵

Bild

Last year, we introduced FlexOlmo, a novel way to train parts of a model independently then combine them later. BAR builds on that idea for a harder problem: how to keep improving a model without having to retrain each time. 🧵

Bild

You can now train, adapt, and eval web agents on your own tasks. We're releasing the full MolmoWeb codebase—the training code, eval harness, annotation tooling, synthetic data pipeline, & client-side code for our demo. 🧵

Bild

Today we're releasing WildDet3D—an open model for monocular 3D object detection in the wild. It works with text, clicks, or 2D boxes, and on zero-shot evals it nearly doubles the best prior scores. 🧵

MolmoBot, our open robotic manipulation suite trained entirely in simulation, now has code, training data, a data generation pipeline, & evals all available. This puts our robotics models within reach of any research lab—no extensive real-world data collection required. 🧵

Bild