Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

AI Research

Research papers, benchmarks, architectures, reasoning and training methods.

#Pallas — 1 article ✕

Google Shows How to Speed Up Video Diffusion Attention on TPUs by 2.4x

Google engineers describe how they implemented sparse spatio-temporal attention for video diffusion models on TPU v6e chips, converting the theoretical sparsity of the Sparse VideoGen (SVG) approach into real hardware speedups. Through a progression of kernel optimizations (full/boundary tile specialization, tile-size tuning, and mask-tile alignment), they reduced attention kernel latency from 96.37ms (naive sparse) to 32.76ms, a 2.40x speedup over dense Splash Attention on a single TPU v6e chip, using 75.6K tokens, 10 heads, and head dimension 128.

Google Developers