Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

All News

TENESYS AI NEWS tracks the most important AI news and explains what changed, why it matters and whether a technology is worth testing.

#Video Diffusion — 1 article ✕

Google Shows How to Speed Up Video Diffusion Attention on TPUs by 2.4x

Google engineers describe how they implemented sparse spatio-temporal attention for video diffusion models on TPU v6e chips, converting the theoretical sparsity of the Sparse VideoGen (SVG) approach into real hardware speedups. Through a progression of kernel optimizations (full/boundary tile specialization, tile-size tuning, and mask-tile alignment), they reduced attention kernel latency from 96.37ms (naive sparse) to 32.76ms, a 2.40x speedup over dense Splash Attention on a single TPU v6e chip, using 75.6K tokens, 10 heads, and head dimension 128.

Google Developers