Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

AI Research

Research papers, benchmarks, architectures, reasoning and training methods.

#JAX — 2 articles ✕

Google's TPU Team Reproduces Ai2's Olmo 3 7B Training Run in MaxText

A Google Cloud TPU engineering team worked with Ai2 to rebuild Olmo 3 7B's pre-training from scratch using MaxText, Google's JAX/XLA training framework, running on TPUs instead of Ai2's original PyTorch/GPU setup. They matched Ai2's published results not just on the training loss curve but on four independent held-out evaluation surfaces across the full ~5.93-trillion-token, 1.41-million-step stage-1 run plus the stage-2 annealing phase. Along the way they ported Olmo 3's unusual architecture (reordered-norm blocks, QK-norm, 3:1 sliding/global attention ratio) into MaxText and caught a data-loader bug that had been quietly inflating apparent performance through memorization rather than genuine learning.

Google Developers

Google Shows How to Speed Up Video Diffusion Attention on TPUs by 2.4x

Google engineers describe how they implemented sparse spatio-temporal attention for video diffusion models on TPU v6e chips, converting the theoretical sparsity of the Sparse VideoGen (SVG) approach into real hardware speedups. Through a progression of kernel optimizations (full/boundary tile specialization, tile-size tuning, and mask-tile alignment), they reduced attention kernel latency from 96.37ms (naive sparse) to 32.76ms, a 2.40x speedup over dense Splash Attention on a single TPU v6e chip, using 75.6K tokens, 10 heads, and head dimension 128.

Google Developers