Google's Antigravity SDK Adds Local Offline Model Support with Gemma 4 26B
In short: Google announced that its Antigravity SDK now supports running agentic workflows locally and offline, with initial support for Gemma 4 26B A4B running via Google AI Edge's LiteRT. Developers can build agents that run entirely on-device, without cloud API costs or internet dependency. Google also demonstrated a hybrid 'Architect-Builder' pattern where a cloud model (Gemini 3.8 Flash) plans tasks and a local swarm of Gemma 4 26B instances does the heavy execution, keeping source code off the cloud.
This summary was generated automatically by AI from Google Developers's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1Antigravity SDK now supports local/offline execution via LiteRT and Gemma 4 26B A4B
- 2Recommended hardware: >24GB VRAM or unified memory
- 3New LiteRTAgentConfig and LocalOpenAIAgentConfig classes for local and OpenAI-compatible backends (Ollama, LM Studio, vLLM)
- 4Hybrid demo: Gemini 3.8 Flash plans tasks (95 cloud tokens) while local Gemma 4 26B swarm executes 97.2% of total tokens (3,322) entirely on-device
- 5Example use case: autonomous security audit/patch workflow across three code modules without uploading source code
| Parameter | Before | Now |
|---|---|---|
| Execution mode | Cloud-only agentic workflows | Local/offline execution via LiteRT with Gemma 4 26B A4B, plus hybrid cloud-local orchestration |
| Hardware requirement | Not specified | Recommended >24GB VRAM or unified memory |
| Cost in hybrid demo | Not specified | 97.2% of tokens (3,322) run locally/offline; only 95 cloud tokens used for Gemini 3.8 Flash planning |
| Compatible backends | Not specified | Supports OpenAI-compatible local servers (Ollama, LM Studio, vLLM) via LocalOpenAIAgentConfig |
Why it matters
This lets agent developers cut API costs, keep sensitive code/data entirely local for compliance reasons, and keep working offline, while still using powerful cloud models for high-level planning when needed. It signals a broader industry shift toward hybrid cloud/edge agent architectures.
What it means for AI agents and contact centers
If your company runs privacy-sensitive contact-center deployments, this pattern is relevant: a lightweight local model could handle on-device tool calling or call-data processing while a cloud model handles complex reasoning, reducing per-call API costs and keeping customer data on-premise. Worth evaluating the hybrid Architect-Builder pattern for latency and cost tradeoffs in voice agent pipelines.
🧪 Worth Testing
Your company could test whether local Gemma 4 26B via LiteRT can handle lightweight agent tasks (e.g., tool calling, call summarization) on-premise to cut cloud costs and address data privacy requirements, while comparing latency and quality against pure cloud-based orchestration.
Antigravity SDK local model support (Gemma 4 26B A4B + LiteRT)· New
Sources
- Google DevelopersOfficialPrimary sourceOriginal article →„Introducing Support for Local AI Models in the Antigravity SDK“23 Sept 2026, 03:00
- Published by source
- 23 Sept 2026, 03:00
- Found by our system
- 2 Oct 2026, 19:51
- Summary generated
- 2 Oct 2026, 19:55
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy