Skip to content
Tuesday, 6 October 2026
Tenesys AI News
Subscribe
Models· Important· 🧪 Worth Testing

Google ships EmbeddingGemma 2, an open multimodal embedding model for RAG

In short: Google has launched EmbeddingGemma 2, a Gemma 4-based embedding model released under Apache 2.0 that places text, code, images, video and audio in one shared 768-dimensional vector space. It's modular, so developers load only the encoders they need: 270M parameters for text and code alone, scaling up to 740M for every modality. Google reports a 14% jump over EmbeddingGemma 1 on the MTEB Code benchmark while keeping the earlier model's multilingual text accuracy.

Source: Google DevelopersGoogleEmbeddingGemma 2Original article ↗

This summary was generated automatically by AI from Google Developers's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1Built on Gemma 4, with modular sizing from 270M (text/code only) to 740M parameters (all modalities)
  • 2Text, code, image, video and audio embeddings share one 768-dimensional vector space
  • 38,192-token context window applies across every modality
  • 4Matryoshka Representation Learning (MRL) lets vectors be truncated from 768 down to 128 dimensions to cut storage
  • 5MTEB Code score up 14% versus EmbeddingGemma 1
  • 6Runs via sentence-transformers, vLLM, Hugging Face Transformers, Ollama, LMStudio and other tooling under Apache 2.0 license
EmbeddingGemma 2
ParameterBeforeNow
Base architectureNot specifiedGemma 4-based decoder
ParametersNot specified270M (text/code) up to 740M (full multimodal)
Embedding spaceNot specifiedShared 768-dimensional space across modalities
Context windowNot specified8,192 tokens
Dimension truncationNot specifiedMatryoshka Representation Learning, 768 down to 128 dimensions
Code retrieval benchmark (MTEB Code)EmbeddingGemma 1 baseline+14% score vs EmbeddingGemma 1
LicenseNot specifiedApache 2.0

Why it matters

Teams building retrieval over mixed content (support tickets, call recordings, screenshots, documentation) can now use a single compact model instead of stitching together separate embedders per modality, simplifying indexing and search pipelines while reducing compute and storage needs.

What it means for AI agents and contact centers

For anyone indexing call transcripts, recorded audio, screenshots of agent screens or support documentation for semantic search, this model offers one shared vector space for all of it with a small memory footprint and configurable dimension size — worth testing against current embedding setups for retrieval accuracy and infrastructure cost, particularly since vector truncation can shrink storage up to 6x.

🧪 Worth Testing

Its ability to combine audio, text and visual content in one vector space at low parameter counts, plus flexible dimension truncation, makes it a strong candidate to benchmark for contact-center search and RAG use cases where storage and latency matter.

EmbeddingGemma 2· New

Sources

  • Google DevelopersOfficialPrimary source
    „EmbeddingGemma 2: The Developer Guide“
    6 Oct 2026, 03:00
    Original article →
  • Google DevelopersOfficial
    „Bring multimodal semantic search to the edge with EmbeddingGemma 2“
    Original article →
Published by source
6 Oct 2026, 03:00
Found by our system
6 Oct 2026, 19:05
Summary generated
6 Oct 2026, 19:07

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy

Google ships EmbeddingGemma 2, an open multimodal embedding model for RAG · TENESYS AI NEWS