Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

Google

Latest news, models and how they changed.

Latest news

Gemini chat adds 'skills' for reusable task instructions, set to replace Gems

Google has begun a global rollout of 'skills' inside the Gemini chat app, a feature that lets the assistant turn recurring chat instructions into reusable shortcuts it can trigger automatically on matching prompts. Users can combine several skills at once to handle more involved tasks, such as layering a tone-of-voice preference together with company-specific formatting rules. Google plans for skills to eventually take over the role currently played by Gems, and existing Gems will be converted into skills automatically once that tool is phased out. For now the feature is live only for users 18 and older, with wider availability to follow.

Gemini App Release Notes

Google's TPU Team Reproduces Ai2's Olmo 3 7B Training Run in MaxText

A Google Cloud TPU engineering team worked with Ai2 to rebuild Olmo 3 7B's pre-training from scratch using MaxText, Google's JAX/XLA training framework, running on TPUs instead of Ai2's original PyTorch/GPU setup. They matched Ai2's published results not just on the training loss curve but on four independent held-out evaluation surfaces across the full ~5.93-trillion-token, 1.41-million-step stage-1 run plus the stage-2 annealing phase. Along the way they ported Olmo 3's unusual architecture (reordered-norm blocks, QK-norm, 3:1 sliding/global attention ratio) into MaxText and caught a data-loader bug that had been quietly inflating apparent performance through memorization rather than genuine learning.

Google Developers

Google Shows How to Speed Up Video Diffusion Attention on TPUs by 2.4x

Google engineers describe how they implemented sparse spatio-temporal attention for video diffusion models on TPU v6e chips, converting the theoretical sparsity of the Sparse VideoGen (SVG) approach into real hardware speedups. Through a progression of kernel optimizations (full/boundary tile specialization, tile-size tuning, and mask-tile alignment), they reduced attention kernel latency from 96.37ms (naive sparse) to 32.76ms, a 2.40x speedup over dense Splash Attention on a single TPU v6e chip, using 75.6K tokens, 10 heads, and head dimension 128.

Google Developers
Agents· Important· 🧪 Worth Testing

Google Cloud Launches Remote MCP Server for gcloud and BigQuery CLI Access

Google Cloud introduced, in public preview, a remote MCP server that exposes the gcloud and bq command-line tools to AI agents via two tools: run_gcloud_command and run_bq_command. The server runs in an isolated, network-restricted sandbox on Google Cloud infrastructure, removing the need for agents to install or maintain local CLI binaries. It uses Agent Identity, OAuth 2.0, and IAM for authentication, integrates with Model Armor to screen for prompt injection, and supports Cloud Audit Logging for full visibility into tool calls.

Google Cloud AI
Agents· Important· 🧪 Worth Testing

Google's Antigravity SDK Adds Local Offline Model Support with Gemma 4 26B

Google announced that its Antigravity SDK now supports running agentic workflows locally and offline, with initial support for Gemma 4 26B A4B running via Google AI Edge's LiteRT. Developers can build agents that run entirely on-device, without cloud API costs or internet dependency. Google also demonstrated a hybrid 'Architect-Builder' pattern where a cloud model (Gemini 3.8 Flash) plans tasks and a local swarm of Gemma 4 26B instances does the heavy execution, keeping source code off the cloud.

Google Developers
Google — AI news · TENESYS AI NEWS