Google Reproduces Ai2's Olmo 3 7B Training in MaxText on TPUs
Google's TPU engineering team reproduced the Allen Institute for AI's fully open Olmo 3 7B model from scratch using MaxText, a JAX/XLA training framework, matching Ai2's PyTorch/GPU reference run on held-out metrics rather than just the training loss curve. The reproduction covered stage-1 pre-training (~5.93T tokens, 1.41M steps) and stage-2 mid-training annealing, including porting Olmo 3's non-standard architecture (reordered-norm block, QK-norm, 3:1 sliding/global attention) to JAX. The team also caught and fixed a subtle data-loader bug that caused training loss to falsely appear better due to memorization, verified by held-out evaluation.
