Google Open-Sources ML Drift, a New GPU Engine for On-Device AI Inference
In short: Google released ML Drift, an open-source, cross-platform GPU compute engine for on-device ML inference, replacing the legacy TFLite GPU delegate. It unifies shader handling across OpenGL ES, OpenCL, Metal and WebGPU, adds 5D tensor support, and introduces stage-aware optimizations for LLM prefill/decode. It is already running in production in Chrome, YouTube Shorts, Google Photos, Meet and partner apps from Adobe and Snap.
This summary was generated automatically by AI from Google Developers's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1Unified tensor virtualization removes the need for separate shader codebases per backend
- 25D tensor support added, removing earlier 4D-only limitation
- 3Stage-aware kernel switching optimizes LLM prefill (compute-bound) vs decode (memory-bound) phases
- 4WebGPU backend via Dawn now runs natively on Windows and Linux, not just in-browser
- 5Up to 12% lower memory overhead vs other frameworks in Gemma benchmarks
- 6Legacy TFLite GPU delegate will no longer receive feature updates
| Parameter | Before | Now |
|---|---|---|
| Tensor dimensionality | 4D tensors only (TFLite GPU delegate) | 5D tensor support out-of-the-box |
| Platform coverage | OpenGL ES, OpenCL, Metal backends separately maintained | Unified shader model via tensor virtualization across OpenGL ES, OpenCL, Metal, WebGPU, plus native Windows/Linux via Da |
| Memory overhead | Baseline (legacy GPU delegate) | Up to 12% lower memory overhead in Gemma benchmarks |
| LLM inference | No stage-aware optimization | Stage-aware kernel switching for prefill vs decode, custom KV cache layout, in-kernel activation quantization |
Why it matters
Teams running speech, vision or generative models on edge devices get a single, open-source runtime instead of juggling fragmented GPU backends, with reported real-world gains like 40% lower frame latency in YouTube Shorts and 30% faster mobile inference in Adobe and Snap apps.
What it means for AI agents and contact centers
If any part of a voice AI pipeline runs on-device (mobile apps, edge boxes, or local LLM-assisted call handling for latency or privacy reasons), this is worth benchmarking against current GPU inference setups, particularly for stage-aware LLM decoding and memory footprint, which affect how many concurrent sessions a device can handle.
🧪 Worth Testing
The stage-aware prefill/decode optimizations and lower memory overhead could directly improve latency and concurrency for any on-device speech or LLM inference used in voice agent deployments.
Sources
- Google DevelopersOfficialPrimary sourceOriginal article →„ML Drift: Next-Gen GPU AI/ML Inference at the Edge“8 Oct 2026, 03:00
- Published by source
- 8 Oct 2026, 03:00
- Found by our system
- 8 Oct 2026, 22:09
- Summary generated
- 8 Oct 2026, 22:10
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy