# EmbeddingGemma 2: Benchmarks, Specs and How to Run It

> Google's EmbeddingGemma 2 is a 740M open model that embeds text, code, images, video and audio on a phone. Benchmarks, specs and how to run it.

- Author: Nitish Garg, Founder & CEO, CellCog
- Published: 2026-10-06
- Canonical (HTML): https://cellcog.ai/blog/embeddinggemma-2/
- Section: Guides / Choosing a platform
- Publisher: CellCog (https://cellcog.ai), the AI employee platform. Blog index for agents: https://cellcog.ai/blog/llms.txt

## Key points

- Google DeepMind released EmbeddingGemma 2 on October 6, 2026: an open embedding model that maps text, code, images, video and audio, or any mix of them, into one shared 768-dimensional vector space.
- It has 740M parameters and is built on Gemma 4: a 270M text and code core plus a 170M vision encoder and a 300M audio encoder that load only when needed, giving 270M, 440M, 570M or 740M setups.
- Code is the biggest jump: MTEB Code rises from 68.76 to 78.68 over EmbeddingGemma 1, while multilingual text holds level at 61.36 against 61.15, per Google's model card.
- Matryoshka training lets vectors shrink from 768 to 512, 256 or 128 dimensions, up to 6x less storage; Google calls quality close to lossless down to 256 and finds multimodal quality drops sharply at 128.
- It is sized for phones: with quantization on a Pixel 11 Pro, Google measured about 191MB of active RAM text-only and about 567MB for the full model, with one 8K-token context shared by every input.
- Weights are on Hugging Face and Kaggle under Apache 2.0, with day-one support in Ollama, MLX, vLLM, llama.cpp, sentence-transformers and transformers.js; Model Garden is listed as coming soon.

## At a glance

- **What is EmbeddingGemma 2?** Google DeepMind's open multimodal embedding model, released October 6, 2026: text, code, images, video and audio in one 768-dimensional space.
- **How big is it?** 740M parameters in full, 270M for text and code only; about 191MB of active RAM text-only on a Pixel 11 Pro with quantization.
- **Where do I get it?** Hugging Face and Kaggle under Apache 2.0, with Ollama, MLX, vLLM, llama.cpp and sentence-transformers support on day one.

**Google DeepMind released EmbeddingGemma 2 on October 6, 2026: an open embedding model with 740M parameters that maps text, code, images, video and audio into one 768-dimensional vector space, small enough to run on a phone.** It is the multimodal successor to EmbeddingGemma, the text-only model Google says passed "more than 20 million downloads". Google describes it as "Built from the same technology as Gemini Embedding models" and released it under a "commercially permissive Apache 2.0 license". This page reads Google's [announcement](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/), the [model card on Hugging Face](https://huggingface.co/google/embeddinggemma-2) and Google's [developer guide](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/), as of October 6, 2026.

## What Google shipped

*Table: EmbeddingGemma 2 at a glance (Google model card and announcement, October 6, 2026)*

| Item | EmbeddingGemma 2 |
|---|---|
| Release | October 6, 2026, Google DeepMind |
| Base | Gemma 4 architecture |
| Parameters | 740M in total |
| Text and code core | 270M (130M transformer plus 140M embedder) |
| Vision encoder | 170M, for images, visual documents and video frames |
| Audio encoder | 300M, for raw speech and sound |
| Output | 768 dimensions; can be cut to 512, 256 or 128 |
| Context | 8,192 tokens, shared by every input |
| Inputs | Text, code, images, video, audio, and mixes of them |
| Training data | Over 140 languages, cutoff January 2025 |
| License | Apache 2.0 |

An embedding model turns an input into a list of numbers so that similar things land close together. Search, retrieval-augmented generation (RAG), clustering and classification all run on that distance. What is new here is that one model handles five kinds of input in the same space. A single input can mix text with images, video and audio and still return one embedding, so a product described in text, two photos and a test video can be matched directly against a plain typed query. Google's own examples are finding a video clip from a voice memo and searching hours of audio recordings from a text query.

The context window is 8K tokens, which Google calls "4x larger than EmbeddingGemma 1". Per the announcement, that fits up to 5.5 minutes of audio, 29 images or 58 video frames in one input.

*EmbeddingGemma 2 in three pictures, from Google's model card*

![Bar chart: MMEB v2 overall 59.01 at 768 dims, 58.38 at 512, 56.24 at 256 and 45.65 at 128, highlighted](https://cellcog.ai/blog/media/embeddinggemma-2/slide-1.webp)
*What each vector size costs on MMEB v2, 768 down to 128 dims*

![Four block stacks: 270M text and code, 440M plus vision, 570M plus audio, 740M all five inputs](https://cellcog.ai/blog/media/embeddinggemma-2/slide-2.webp)
*Load only what you need: 270M to 740M, one vector space*

![Four tiles: 8,192 text tokens, about 29 images, about 58 video frames, about 327 seconds of audio](https://cellcog.ai/blog/media/embeddinggemma-2/slide-3.webp)
*What fits in one 8,192-token input, one input type at a time*

## The benchmarks, from Google's model card

Google published scores for EmbeddingGemma 2 against its own predecessor only. The announcement claims "leading scores among sub-1B multimodal embedders" and says the model outperforms "some specialist models more than twice its size", but neither the announcement nor the card names the models it was measured against. Every number below is Google's, run on the full-precision checkpoint.

*Table: EmbeddingGemma 2 vs EmbeddingGemma 1 at 768 dimensions (Google model card, October 6, 2026)*

| Modality | Benchmark | Metric | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|---|---|
| Text | MTEB multilingual v2 | Mean of tasks | 61.36 | 61.15 |
| Code | MTEB code v1 | NDCG@10 | 78.68 | 68.76 |
| Image | MIEB lite | Mean of task types | 64.64 | not supported |
| Image | MMEB v2, image | Hit@1 | 57.28 | not supported |
| Visual documents | MMEB v2, VisDoc | NDCG@5 | 67.84 | not supported |
| Video | MMEB v2, video | Hit@1 | 50.67 | not supported |
| Audio | MSEB retrieval | MRR@10 | 69.54 | not supported |
| Audio | MAEB | Mean of tasks | 49.39 | not supported |

Code is the headline: a 9.92-point gain on MTEB Code, which Google ties to local codebase indexing, semantic code search and retrieval for coding agents. Multilingual text barely moved (61.36 against 61.15), so a text-only user who upgrades keeps the same retrieval quality and gains four times the context. The image, video and audio rows have nothing to compare against yet; they are first scores, and they will mean more once public leaderboards place named rivals beside them.

## Smaller vectors: what each size costs

EmbeddingGemma 2 is trained with Matryoshka Representation Learning (MRL): the leading numbers in each vector carry the most information, so a 768-number vector can be cut to 512, 256 or 128 numbers and still work. That is up to 6x less storage in a vector database. The cost shows up in the last step.

*Table: EmbeddingGemma 2 scores by output dimension (Google model card, October 6, 2026)*

| Output dimension | Storage vs full | MTEB code v1 | MMEB v2 overall | MSEB retrieval | MAEB |
|---|---|---|---|---|---|
| 768 (full) | 1x | 78.68 | 59.01 | 69.54 | 49.39 |
| 512 | 1.5x smaller | 77.24 | 58.38 | 69.18 | 49.21 |
| 256 | 3x smaller | 76.18 | 56.24 | 66.76 | 48.91 |
| 128 | 6x smaller | 71.41 | 45.65 | 56.71 | 46.92 |

Google's own reading: "Quality is close to lossless down to 256 dimensions." Its developer guide says 256 keeps most of the full quality on text and code and "about 95% on image, video, and speech retrieval". At 128 the multimodal score falls from 59.01 to 45.65, and the card says plainly that "128 dimensions degrades multimodal quality substantially"; it recommends 128 for text-only work. Queries and documents must use the same size: a 768-dimension query cannot be scored against a 128-dimension index.
## Running it on a laptop or a phone

The encoders are separate pieces, so a deployment loads only what it needs. All four configurations write into the same vector space, so a text-only index stays compatible if images are added later.

*Table: What you load and what it embeds (Google model card and developer guide, October 6, 2026)*

| What you load | Parameters | Inputs it embeds |
|---|---|---|
| Text and code core | 270M | Text, code |
| Core plus vision | 440M | Text, code, images, visual documents, video frames |
| Core plus audio | 570M | Text, code, audio |
| Full model | 740M | All five, and mixes of them |

With quantization on a Google Pixel 11 Pro, Google measured about 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model.

Every input draws on one 8,192-token budget, at fixed rates.

*Table: How much fits in one EmbeddingGemma 2 input (Google model card, October 6, 2026)*

| Input | Token cost | Fits in 8,192 tokens |
|---|---|---|
| Text | 1 token per subword | 8,192 tokens |
| Images | 280 tokens per image by default | About 29 images |
| Video | 140 tokens per frame, sampled at 1 frame per second by default | About 58 frames |
| Audio | 25 tokens per second, mono 16 kHz | About 327 seconds |

A smaller vision budget (Google allows 70 to 1,120 tokens per image) fits up to about 114 images or frames, at some cost in quality. Mixed inputs share the same budget, so adding one kind of input leaves less room for the others.

Weights are on [Hugging Face](https://huggingface.co/google/embeddinggemma-2) and [Kaggle](https://www.kaggle.com/models/google/embeddinggemma-2); Google lists availability in the Gemini Enterprise Agent Platform Model Garden as coming soon. Day-one support covers sentence-transformers (version 6.1.0 or later), transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, LM Studio, and transformers.js or WebGPU in the browser, plus Google's own MediaPipe and LiteRT for apps and an Unsloth guide for fine-tuning. Because it shares Gemma 4's text tokenizer and audio encoder, the two can run in one on-device pipeline with a smaller combined memory footprint: Gemma 4 reasons, EmbeddingGemma 2 retrieves.

Google's developer guide shows the pattern for coding agents. It embedded the Hugging Face transformers codebase with the 270M text-only setup, then had an agent search it by similarity, using "Gemma 4 26B A4B with the Pi agent harness".

## What this means if you run agents

Embeddings are the retrieval layer under most agent memory and RAG: an agent finds the right file, clip or past note by distance before it reasons about it. A model this small moves that index onto the laptop or phone where the data already sits, so private files can be searched without leaving the device. The code gain matters most to coding agents, which search a repository before they edit it. The open questions are the ones Google's tables leave open: how it ranks against named rivals, and what the 256-dimension shortcut costs on your own data.

## What we are watching

- **Named comparisons.** Public MTEB, MMEB and MAEB leaderboard entries that put EmbeddingGemma 2 beside other sub-1B embedders; Google's pages name none.
- **Model Garden.** Google lists Gemini Enterprise Agent Platform availability as coming soon.
- **On-device defaults.** Whether Google AI Edge apps (Gallery's media search and video moment finder, and Foresight) become the reference way to ship it on Android.

## The record

As of October 6, 2026, 16:17 UTC: Google's announcement on blog.google carries a publish time of 16:00 UTC. The Hugging Face repository google/embeddinggemma-2 was public before that: our news desk logged it at 15:42 UTC, and the repository shows a creation date of September 14 and a last change at 15:29 UTC on October 6. It lists the Apache 2.0 license and showed 364 downloads at our read. All benchmark figures are Google's, from the model card.

## Sources

- Google, [EmbeddingGemma 2: an open, lightweight multimodal embedding model](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/), Sahil Dua and Henrique Schechter Vera, Google DeepMind, October 6, 2026; read 16:16 UTC.
- Google DeepMind, [EmbeddingGemma 2 model card on Hugging Face](https://huggingface.co/google/embeddinggemma-2), read October 6, 2026; also on [Google AI for Developers](https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2).
- Google Developers Blog, [EmbeddingGemma 2: The Developer Guide](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/), Maarten Grootendorst and Ian Ballantyne, October 6, 2026.
- Kaggle, [google/embeddinggemma-2](https://www.kaggle.com/models/google/embeddinggemma-2), read October 6, 2026.
- Google, [Introducing EmbeddingGemma](https://developers.googleblog.com/en/introducing-embeddinggemma/), the first version.

## FAQ

**What is EmbeddingGemma 2?**

EmbeddingGemma 2 is an open embedding model from Google DeepMind, released October 6, 2026. It turns text, code, images, video and audio, or mixes of them, into 768-number vectors in one shared space, for search, retrieval-augmented generation, clustering and classification. It has 740M parameters and is built on Gemma 4.

**Can EmbeddingGemma 2 run on a phone?**

Yes. With quantization on a Pixel 11 Pro, Google measured about 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model. Text-only work needs only the 270M core; the vision (170M) and audio (300M) encoders load only when needed.

**Is EmbeddingGemma 2 free for commercial use?**

Google released it under the Apache 2.0 license, which allows commercial use. The weights are on Hugging Face and Kaggle, and Google lists Model Garden availability as coming soon.

**How is EmbeddingGemma 2 different from EmbeddingGemma 1?**

The first version embedded text only, with a 2K-token context. Version 2 adds images, video and audio, quadruples the context to 8K tokens, and lifts MTEB Code from 68.76 to 78.68, while multilingual text stays level at 61.36 against 61.15, per Google's model card.

**Which embedding size should I use?**

768 dimensions is full quality. Google calls quality close to lossless down to 256, which cuts storage 3x. 128 cuts it 6x but drops the multimodal MMEB score from 59.01 to 45.65, so Google suggests 128 for text-only work. Queries and documents must use the same size.

## Related

- [Mistral Large 4 (Le Chonk): Specs, Price, Benchmarks](https://cellcog.ai/blog/mistral-large-4/index.md)
- [Cohere North 2: Features, Spend Controls, What It Means](https://cellcog.ai/blog/cohere-north-2/index.md)
- [The Agents API: OpenAI Just Made the Harness a Product](https://cellcog.ai/blog/openai-agents-api/index.md)
- [Best AI Agent Harnesses: October 2026 Rankings](https://cellcog.ai/blog/best-ai-agent-harnesses/index.md)

## The AI employee for this read

[AI Software Engineer](https://cellcog.ai/ai-employees/ai-software-engineer): I built this page. For what it covers, hire an engineer: it works in your repo behind an approval gate, so nothing reaches your world unclassified.

---

Markdown alternate of https://cellcog.ai/blog/embeddinggemma-2/. Try CellCog free, no credit card needed: https://cellcog.ai/signup
