Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

EmbeddingGemma 2: Benchmarks, Specs and How to Run It

At a glanceQuick answers
What is EmbeddingGemma 2?
Google DeepMind’s open multimodal embedding model, released October 6, 2026: text, code, images, video and audio in one 768-dimensional space.
How big is it?
740M parameters in full, 270M for text and code only; about 191MB of active RAM text-only on a Pixel 11 Pro with quantization.
Where do I get it?
Hugging Face and Kaggle under Apache 2.0, with Ollama, MLX, vLLM, llama.cpp and sentence-transformers support on day one.
Data illustration on off-white paper: five teal streams labeled text, code, image, video and audio flow into a smartphone that emits an amber grid of dots, beside the numbers 740M parameters, 768 dimensions and 8K context
Fig 0Five kinds of input, one vector space, one phone. Made by CellCog's image agent, running GPT Image 2.5.

Google DeepMind released EmbeddingGemma 2 on October 6, 2026: an open embedding model with 740M parameters that maps text, code, images, video and audio into one 768-dimensional vector space, small enough to run on a phone. It is the multimodal successor to EmbeddingGemma, the text-only model Google says passed “more than 20 million downloads”. Google describes it as “Built from the same technology as Gemini Embedding models” and released it under a “commercially permissive Apache 2.0 license”. This page reads Google’s announcement, the model card on Hugging Face and Google’s developer guide, as of October 6, 2026.

On this page · 8 sectionsOpen
  1. What Google shipped
  2. The benchmarks, from Google’s model card
  3. Smaller vectors: what each size costs
  4. Running it on a laptop or a phone
  5. What this means if you run agents
  6. What we are watching
  7. The record
  8. Sources
Key points6 · 9 min full read
  1. Five small input icons flowing into one cluster of dots: many kinds of input, one vector space.
    Google DeepMind released EmbeddingGemma 2 on October 6, 2026: an open embedding model that maps text, code, images, video and audio, or any mix of them, into one shared 768-dimensional vector space.
  2. Three stacked blocks of different sizes: modules that load only when needed.
    It has 740M parameters and is built on Gemma 4: a 270M text and code core plus a 170M vision encoder and a 300M audio encoder that load only when needed, giving 270M, 440M, 570M or 740M setups.
  3. Code braces with an arrow pointing up: better code retrieval.
    Code is the biggest jump: MTEB Code rises from 68.76 to 78.68 over EmbeddingGemma 1, while multilingual text holds level at 61.36 against 61.15, per Google’s model card.
  4. Nested squares from large to small: vectors that can be cut shorter.
    Matryoshka training lets vectors shrink from 768 to 512, 256 or 128 dimensions, up to 6x less storage; Google calls quality close to lossless down to 256 and finds multimodal quality drops sharply at 128.
  5. A phone with a chip on its screen: the model running on the device.
    It is sized for phones: with quantization on a Pixel 11 Pro, Google measured about 191MB of active RAM text-only and about 567MB for the full model, with one 8K-token context shared by every input.
  6. An arrow dropping into an open box: a free download.
    Weights are on Hugging Face and Kaggle under Apache 2.0, with day-one support in Ollama, MLX, vLLM, llama.cpp, sentence-transformers and transformers.js; Model Garden is listed as coming soon.

§ 01What Google shipped

Item EmbeddingGemma 2
Release October 6, 2026, Google DeepMind
Base Gemma 4 architecture
Parameters 740M in total
Text and code core 270M (130M transformer plus 140M embedder)
Vision encoder 170M, for images, visual documents and video frames
Audio encoder 300M, for raw speech and sound
Output 768 dimensions; can be cut to 512, 256 or 128
Context 8,192 tokens, shared by every input
Inputs Text, code, images, video, audio, and mixes of them
Training data Over 140 languages, cutoff January 2025
License Apache 2.0
Table 1EmbeddingGemma 2 at a glance (Google model card and announcement, October 6, 2026)

An embedding model turns an input into a list of numbers so that similar things land close together. Search, retrieval-augmented generation (RAG), clustering and classification all run on that distance. What is new here is that one model handles five kinds of input in the same space. A single input can mix text with images, video and audio and still return one embedding, so a product described in text, two photos and a test video can be matched directly against a plain typed query. Google’s own examples are finding a video clip from a voice memo and searching hours of audio recordings from a text query.

The context window is 8K tokens, which Google calls “4x larger than EmbeddingGemma 1”. Per the announcement, that fits up to 5.5 minutes of audio, 29 images or 58 video frames in one input.

§ 02The benchmarks, from Google’s model card

Google published scores for EmbeddingGemma 2 against its own predecessor only. The announcement claims “leading scores among sub-1B multimodal embedders” and says the model outperforms “some specialist models more than twice its size”, but neither the announcement nor the card names the models it was measured against. Every number below is Google’s, run on the full-precision checkpoint.

Modality Benchmark Metric EmbeddingGemma 2 EmbeddingGemma 1
Text MTEB multilingual v2 Mean of tasks 61.36 61.15
Code MTEB code v1 NDCG@10 78.68 68.76
Image MIEB lite Mean of task types 64.64 not supported
Image MMEB v2, image Hit@1 57.28 not supported
Visual documents MMEB v2, VisDoc NDCG@5 67.84 not supported
Video MMEB v2, video Hit@1 50.67 not supported
Audio MSEB retrieval MRR@10 69.54 not supported
Audio MAEB Mean of tasks 49.39 not supported
Table 2EmbeddingGemma 2 vs EmbeddingGemma 1 at 768 dimensions (Google model card, October 6, 2026)
MTEB Code v1 score, EmbeddingGemma 1 vs 2, Google's model cardBar chart of MTEB Code v1 scores from Google's model card: EmbeddingGemma 1 at 68.76 and EmbeddingGemma 2 highlighted at 78.68EmbeddingGemma 168.76EmbeddingGemma 278.68MTEB Code v1 score, EmbeddingGemma 1 vs 2, Google's model cardBar chart of MTEB Code v1 scores from Google's model card: EmbeddingGemma 1 at 68.76 and EmbeddingGemma 2 highlighted at 78.68EmbeddingGemma 168.76EmbeddingGemma 278.68
Fig 1MTEB Code v1 score, EmbeddingGemma 1 vs 2, Google's model card

Code is the headline: a 9.92-point gain on MTEB Code, which Google ties to local codebase indexing, semantic code search and retrieval for coding agents. Multilingual text barely moved (61.36 against 61.15), so a text-only user who upgrades keeps the same retrieval quality and gains four times the context. The image, video and audio rows have nothing to compare against yet; they are first scores, and they will mean more once public leaderboards place named rivals beside them.

§ 03Smaller vectors: what each size costs

EmbeddingGemma 2 is trained with Matryoshka Representation Learning (MRL): the leading numbers in each vector carry the most information, so a 768-number vector can be cut to 512, 256 or 128 numbers and still work. That is up to 6x less storage in a vector database. The cost shows up in the last step.

Output dimension Storage vs full MTEB code v1 MMEB v2 overall MSEB retrieval MAEB
768 (full) 1x 78.68 59.01 69.54 49.39
512 1.5x smaller 77.24 58.38 69.18 49.21
256 3x smaller 76.18 56.24 66.76 48.91
128 6x smaller 71.41 45.65 56.71 46.92
Scroll to compare all columns
Table 3EmbeddingGemma 2 scores by output dimension (Google model card, October 6, 2026)
MMEB v2 overall score by output dimension, EmbeddingGemma 2 model cardBar chart of the MMEB v2 overall multimodal score by output dimension: 768 dims 59.01, 512 dims 58.38, 256 dims 56.24, and 128 dims highlighted at 45.65768 dims59.01512 dims58.38256 dims56.24128 dims45.65MMEB v2 overall score by output dimension, EmbeddingGemma 2 model cardBar chart of the MMEB v2 overall multimodal score by output dimension: 768 dims 59.01, 512 dims 58.38, 256 dims 56.24, and 128 dims highlighted at 45.65768 dims59.01512 dims58.38256 dims56.24128 dims45.65
Fig 2MMEB v2 overall score by output dimension, EmbeddingGemma 2 model card

Google’s own reading: “Quality is close to lossless down to 256 dimensions.” Its developer guide says 256 keeps most of the full quality on text and code and “about 95% on image, video, and speech retrieval”. At 128 the multimodal score falls from 59.01 to 45.65, and the card says plainly that “128 dimensions degrades multimodal quality substantially”; it recommends 128 for text-only work. Queries and documents must use the same size: a 768-dimension query cannot be scored against a 128-dimension index.

§ 04Running it on a laptop or a phone

The encoders are separate pieces, so a deployment loads only what it needs. All four configurations write into the same vector space, so a text-only index stays compatible if images are added later.

What you load Parameters Inputs it embeds
Text and code core 270M Text, code
Core plus vision 440M Text, code, images, visual documents, video frames
Core plus audio 570M Text, code, audio
Full model 740M All five, and mixes of them
Table 4What you load and what it embeds (Google model card and developer guide, October 6, 2026)

With quantization on a Google Pixel 11 Pro, Google measured about 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model.

Every input draws on one 8,192-token budget, at fixed rates.

Input Token cost Fits in 8,192 tokens
Text 1 token per subword 8,192 tokens
Images 280 tokens per image by default About 29 images
Video 140 tokens per frame, sampled at 1 frame per second by default About 58 frames
Audio 25 tokens per second, mono 16 kHz About 327 seconds
Table 5How much fits in one EmbeddingGemma 2 input (Google model card, October 6, 2026)

A smaller vision budget (Google allows 70 to 1,120 tokens per image) fits up to about 114 images or frames, at some cost in quality. Mixed inputs share the same budget, so adding one kind of input leaves less room for the others.

Weights are on Hugging Face and Kaggle; Google lists availability in the Gemini Enterprise Agent Platform Model Garden as coming soon. Day-one support covers sentence-transformers (version 6.1.0 or later), transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, LM Studio, and transformers.js or WebGPU in the browser, plus Google’s own MediaPipe and LiteRT for apps and an Unsloth guide for fine-tuning. Because it shares Gemma 4’s text tokenizer and audio encoder, the two can run in one on-device pipeline with a smaller combined memory footprint: Gemma 4 reasons, EmbeddingGemma 2 retrieves.

Google’s developer guide shows the pattern for coding agents. It embedded the Hugging Face transformers codebase with the 270M text-only setup, then had an agent search it by similarity, using “Gemma 4 26B A4B with the Pi agent harness”.

§ 05What this means if you run agents

Embeddings are the retrieval layer under most agent memory and RAG: an agent finds the right file, clip or past note by distance before it reasons about it. A model this small moves that index onto the laptop or phone where the data already sits, so private files can be searched without leaving the device. The code gain matters most to coding agents, which search a repository before they edit it. The open questions are the ones Google’s tables leave open: how it ranks against named rivals, and what the 256-dimension shortcut costs on your own data.

§ 06What we are watching

  • Named comparisons. Public MTEB, MMEB and MAEB leaderboard entries that put EmbeddingGemma 2 beside other sub-1B embedders; Google’s pages name none.
  • Model Garden. Google lists Gemini Enterprise Agent Platform availability as coming soon.
  • On-device defaults. Whether Google AI Edge apps (Gallery’s media search and video moment finder, and Foresight) become the reference way to ship it on Android.

§ 07The record

As of October 6, 2026, 16:17 UTC: Google’s announcement on blog.google carries a publish time of 16:00 UTC. The Hugging Face repository google/embeddinggemma-2 was public before that: our news desk logged it at 15:42 UTC, and the repository shows a creation date of September 14 and a last change at 15:29 UTC on October 6. It lists the Apache 2.0 license and showed 364 downloads at our read. All benchmark figures are Google’s, from the model card.

§ 08Sources

Frequently asked5 questions

Q1What is EmbeddingGemma 2?

EmbeddingGemma 2 is an open embedding model from Google DeepMind, released October 6, 2026. It turns text, code, images, video and audio, or mixes of them, into 768-number vectors in one shared space, for search, retrieval-augmented generation, clustering and classification. It has 740M parameters and is built on Gemma 4.

Q2Can EmbeddingGemma 2 run on a phone?

Yes. With quantization on a Pixel 11 Pro, Google measured about 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model. Text-only work needs only the 270M core; the vision (170M) and audio (300M) encoders load only when needed.

Q3Is EmbeddingGemma 2 free for commercial use?

Google released it under the Apache 2.0 license, which allows commercial use. The weights are on Hugging Face and Kaggle, and Google lists Model Garden availability as coming soon.

Q4How is EmbeddingGemma 2 different from EmbeddingGemma 1?

The first version embedded text only, with a 2K-token context. Version 2 adds images, video and audio, quadruples the context to 8K tokens, and lifts MTEB Code from 68.76 to 78.68, while multilingual text stays level at 61.36 against 61.15, per Google’s model card.

Q5Which embedding size should I use?

768 dimensions is full quality. Google calls quality close to lossless down to 256, which cuts storage 3x. 128 cuts it 6x but drops the multimodal MMEB score from 59.01 to 45.65, so Google suggests 128 for text-only work. Queries and documents must use the same size.

Published 06 October 2026 All Choosing a platform →