during in Olmo 3 we thought long context is just finding good data nope! model architecture matters & it's hard to recover if mess it up led by @abertsch.bsky.social, we release many pretrain runs w/ small arch changes and show huge long context performance diffs
Recipes for teaching language models to handle long inputs don't work equally well across model families. We wanted to know why—is it the architecture, the training data, or both? 🧵