Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

Olmo 3 7B

News

Google's TPU Team Reproduces Ai2's Olmo 3 7B Training Run in MaxText

A Google Cloud TPU engineering team worked with Ai2 to rebuild Olmo 3 7B's pre-training from scratch using MaxText, Google's JAX/XLA training framework, running on TPUs instead of Ai2's original PyTorch/GPU setup. They matched Ai2's published results not just on the training loss curve but on four independent held-out evaluation surfaces across the full ~5.93-trillion-token, 1.41-million-step stage-1 run plus the stage-2 annealing phase. Along the way they ported Olmo 3's unusual architecture (reordered-norm blocks, QK-norm, 3:1 sliding/global attention ratio) into MaxText and caught a data-loader bug that had been quietly inflating apparent performance through memorization rather than genuine learning.

Google Developers
Olmo 3 7B · TENESYS AI NEWS