Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

Google Shows How to Speed Up Video Diffusion Attention on TPUs by 2.4x

In short: Google engineers describe how they implemented sparse spatio-temporal attention for video diffusion models on TPU v6e chips, converting the theoretical sparsity of the Sparse VideoGen (SVG) approach into real hardware speedups. Through a progression of kernel optimizations (full/boundary tile specialization, tile-size tuning, and mask-tile alignment), they reduced attention kernel latency from 96.37ms (naive sparse) to 32.76ms, a 2.40x speedup over dense Splash Attention on a single TPU v6e chip, using 75.6K tokens, 10 heads, and head dimension 128.

Source: Google DevelopersGoogleOriginal article ↗

This summary was generated automatically by AI from Google Developers's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1Naive sparse attention (B1) was actually 22% slower than dense attention despite skipping tiles, due to unnecessary intra-tile masking
  • 2Full/boundary tile specialization (B2) cut latency from 96.37ms to 54.12ms (31% faster than dense)
  • 3Tile-size tuning showed the configuration with least masking work was not the fastest
  • 4Final tile-aligned sparse traversal (B3) reduced latency to 32.76ms, a 2.40x speedup over dense Splash Attention
  • 5Token permutation from frame-major to temporal-major layout improved temporal-head attention performance (44.40ms vs 74.14ms without permutation)

Why it matters

Faster video generation inference on TPUs matters for companies building video diffusion products, showing that theoretical sparsity gains require careful hardware-aware kernel engineering to realize actual speedups — a lesson applicable broadly to optimizing any sparse/attention-based inference workload.

Sources

  • Google DevelopersOfficialPrimary source
    „Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs“
    30 Sept 2026, 03:00
    Original article →
Published by source
30 Sept 2026, 03:00
Found by our system
2 Oct 2026, 19:51
Summary generated
2 Oct 2026, 20:05

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy