Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe
Agents· Important· 🧪 Worth Testing

Google's Antigravity SDK Adds Local Offline Model Support with Gemma 4 26B

In short: Google announced that its Antigravity SDK now supports running agentic workflows locally and offline, with initial support for Gemma 4 26B A4B running via Google AI Edge's LiteRT. Developers can build agents that run entirely on-device, without cloud API costs or internet dependency. Google also demonstrated a hybrid 'Architect-Builder' pattern where a cloud model (Gemini 3.8 Flash) plans tasks and a local swarm of Gemma 4 26B instances does the heavy execution, keeping source code off the cloud.

Source: Google DevelopersGoogleGemma 4 26B A4BOriginal article ↗

This summary was generated automatically by AI from Google Developers's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1Antigravity SDK now supports local/offline execution via LiteRT and Gemma 4 26B A4B
  • 2Recommended hardware: >24GB VRAM or unified memory
  • 3New LiteRTAgentConfig and LocalOpenAIAgentConfig classes for local and OpenAI-compatible backends (Ollama, LM Studio, vLLM)
  • 4Hybrid demo: Gemini 3.8 Flash plans tasks (95 cloud tokens) while local Gemma 4 26B swarm executes 97.2% of total tokens (3,322) entirely on-device
  • 5Example use case: autonomous security audit/patch workflow across three code modules without uploading source code
Gemma 4 26B A4B
ParameterBeforeNow
Execution modeCloud-only agentic workflowsLocal/offline execution via LiteRT with Gemma 4 26B A4B, plus hybrid cloud-local orchestration
Hardware requirementNot specifiedRecommended >24GB VRAM or unified memory
Cost in hybrid demoNot specified97.2% of tokens (3,322) run locally/offline; only 95 cloud tokens used for Gemini 3.8 Flash planning
Compatible backendsNot specifiedSupports OpenAI-compatible local servers (Ollama, LM Studio, vLLM) via LocalOpenAIAgentConfig

Why it matters

This lets agent developers cut API costs, keep sensitive code/data entirely local for compliance reasons, and keep working offline, while still using powerful cloud models for high-level planning when needed. It signals a broader industry shift toward hybrid cloud/edge agent architectures.

What it means for AI agents and contact centers

If your company runs privacy-sensitive contact-center deployments, this pattern is relevant: a lightweight local model could handle on-device tool calling or call-data processing while a cloud model handles complex reasoning, reducing per-call API costs and keeping customer data on-premise. Worth evaluating the hybrid Architect-Builder pattern for latency and cost tradeoffs in voice agent pipelines.

🧪 Worth Testing

Your company could test whether local Gemma 4 26B via LiteRT can handle lightweight agent tasks (e.g., tool calling, call summarization) on-premise to cut cloud costs and address data privacy requirements, while comparing latency and quality against pure cloud-based orchestration.

Antigravity SDK local model support (Gemma 4 26B A4B + LiteRT)· New

Sources

  • Google DevelopersOfficialPrimary source
    „Introducing Support for Local AI Models in the Antigravity SDK“
    23 Sept 2026, 03:00
    Original article →
Published by source
23 Sept 2026, 03:00
Found by our system
2 Oct 2026, 19:51
Summary generated
2 Oct 2026, 19:55

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy