Ai2

@ai2.bsky.social

Breakthrough AI to solve the world's biggest problems. › Join us: http://allenai.org/careers › Get our newsletter: https://share.hsforms.com/1uJkWs5aDRHWhiky3aHooIg3ioxm

Super excited to share that our method to convert language models to byte-level has been published in @nature.com! This was originally the approach behind Bolmo. We now generalized it to other model families. All checkpoints (including new Qwen and Llama models) and data are public!

Ai2@ai2.bsky.social · yesterday

Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff.ly/l5ILoNT

Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff.ly/l5ILoNT

Bild

We're heading to #COLM2026 next week! Four days of workshops, posters, & talks on our latest fully open AI research, from Olmo Hybrid to evals for AI-assisted scientific writing. 🧵

BildBildBildBild

Introducing AstaBrief 8B, an open model that turns complex research questions + literature excerpts into cited reports. Run it locally on your own hardware, with open weights + training data you can inspect & build on. 🧵 🤗 Download: buff.ly/HkIcmfR

We’re releasing Olmo-core 3—open training infrastructure for large mixture-of-experts (MoE) models. It’s a core system behind the next generation of Olmo, designed to scale into the trillion-parameter range. 🧵 💻 GitHub repo: buff.ly/CGnP8ax

Bild

Google Cloud put Olmo 3’s reproducibility to the test, rerunning our 7B pretraining & mid-training on Cloud TPUs and matching our original run on held-out evals. Reproducibility matters for science + trustworthy AI. That’s what fully open makes possible. 🤝 developers.googleblog.com/reproducing-...

Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs- Google Developers Blog

Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.

developers.googleblog.com

Can a fully open model help make AI more “prosocial” through a game? Soham Padia built Steering Arena, where players try to elicit kind & respectful responses from Olmo 3. Surprisingly, strings like “Undert! AH :-) Rog Appl)” were highly effective. 👇 buff.ly/adP47yO

Bild

How should future scientists learn to interrogate AI tools for discovery? Read how UW students put AutoDiscovery to the test as part of an academic challenge earlier this year, then try the tool for yourself—credits are now extended through Dec. 31. 🧵 buff.ly/fXoQxxY

Bild

Training an LLM by showing it answers people prefer – and ones they don’t – can improve it while quietly worsening behaviors. Goodfire AI used Ai2’s open post-training stack to predict how a full training run would change responses to different prompts. 🧵 buff.ly/RuJxLM1

Bild

ML emulators mimic climate processes faster than traditional models. Next is coupling atmosphere & ocean emulators so their predictions feed into each other as the simulation runs. With the E3SM team, we built a system that does that: SamudrACE-E3SMv3. 🧵 buff.ly/9ShSTCM

BildBild

Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions! See 🧵

Ai2@ai2.bsky.social · last mo.

Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 buff.ly/bTcvqJf

Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 buff.ly/bTcvqJf

Bild

At an event on August 27, we brought together AI researchers, scientists, & medical practitioners to explore what AI needs to do better to meaningfully advance science. Five ideas kept coming up. 🧵

Bild

AI’s biggest role in science may not be answering questions. It may be helping scientists find which questions are worth asking. At Providence Swedish, AutoDiscovery surfaced an unexpected immune signal in a heavily studied cancer dataset—and follow-up research confirmed it. 🧵 buff.ly/LTacUQl

A Thai research team adapted our Dolma data-curation toolkit to build Mangosteen, a 47B-token corpus for Thai LLMs. They used Dolma to filter widely used web datasets into a smaller corpus that improved Thai LLM performance despite using less data. 🧵 buff.ly/fvLSIru

Bild

What kinds of training data shape different model capabilities? A Georgia Tech team used our fully open model flow to trace Olmo’s performance on social/general reasoning and social-science/STEM knowledge tests back to the types of text it trained on. 🧵 buff.ly/cQEzbUn

Bild

Today we're introducing a preview of TutorMoments, a framework that measures whether AI tutors can make one of the hardest calls in teaching: when to step in and help a student, & when to hold back and let them do the heavy thinking. 🧵

Bild

We're expanding our partnership with @hf.co to accelerate open science. Our storage on the Hub is roughly tripling to ~2 petabytes, & our downloads now run at high speed—even for our largest datasets & multi-checkpoint models. 🧵

Bild

📸 140+ people joined us during #SeattleTechWeek to learn how AI gets built at Ai2. Research lead Iz Beltagy walked through what it takes to train a truly open LLM—from large-scale training runs to eval before release. Thanks to everyone who came + asked thoughtful questions. 🤝

Bild

When a model writes, where do its words come from? Are they new, or do they match exactly with language it saw in training? An AI-writing detector can't tell you. @tuhinchakr.bsky.social's group at Stony Brook has been dissecting AI-generated prose with our infini-gram engine. 🧵

Bild

The organizations best positioned to use Earth-observation models – those working in conservation, food security, & disaster response – often can't run these models at the scale they need. That's an infrastructure problem. So we built the OlmoEarth Platform to solve it. 🧵

Bild

As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.

As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.

Bild

As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.

Bild