October 6, 2026 — Google DeepMind released EmbeddingGemma 2, an open-weight embedding model that places text, code, images, video and audio into one shared 768-dimension vector space. It is built on Gemma 4, ships under Apache 2.0, and comes in sizes from 270M to 740M parameters so it can run on a phone or in a browser. If you build search or retrieval-augmented generation (RAG) and you have been stitching together a text embedder, an image embedder and a speech pipeline, this is a direct attempt to replace the stack with one small model.

This post covers what Google says is in the release, how the modular design works, what the numbers do and do not tell you, and a practical way to try it. Everything below comes from Google's own announcement and developer guide unless marked otherwise.
What EmbeddingGemma 2 is, in plain language
An embedding model turns content into a list of numbers (a vector) so that similar things end up close together. If you want the basics first, our explainer on what an embedding is walks through the idea with examples. Search, recommendations, deduplication and RAG all lean on this one primitive: embed your documents once, embed each query at run time, return the nearest vectors.
Until recently, most open embedders were text-only. If you wanted to search images or audio, you ran a second or third model, and the vectors from different models could not be compared with each other. Google's pitch for EmbeddingGemma 2 is that all inputs pass through a shared backbone and land in the same dimensional space, so a text query like "dog barking at a mailman" can pull back a photo, a video frame or an audio clip directly. That is called cross-modal retrieval, and it is the headline capability.
Per Google's announcement and the developer guide, the key facts are:
| Item | What Google states |
|---|---|
| Base | Gemma 4 |
| License | Apache 2.0 (commercial use allowed) |
| Modalities | Text, code, images, video, audio |
| Output dimension | 768, reducible to 512, 256 or 128 via Matryoshka truncation |
| Context window | 8,192 tokens shared across modalities (4x the first EmbeddingGemma) |
| Code benchmark | 78.68 on MTEB Code, up 9.92 points over v1 |
| Audio | Leading scores among sub-1B models on MAEB, per Google |
| Weights | Hugging Face and Kaggle |
The modular design: load only what you need
The most practical design choice is that the model is not one indivisible blob. The developer guide lists configurations:
- 270M for text and code only
- 440M for text plus vision
- 570M for text plus audio
- 740M for the full multimodal model
Google's announcement page also mentions standalone vision and audio encoders in the low hundreds of millions of parameters. The exact split matters less than the principle: if your product only needs code search, you load the 270M piece and skip the rest. That is what makes the on-device story credible. Google reports about 191MB of active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Google Pixel 11 Pro. Those are vendor-reported figures on one device, so treat them as a ceiling for planning, not a guarantee for your hardware.
The 8K context window is shared across modalities. Google translates that into concrete limits for local hardware: about 5.5 minutes of audio, 29 images, or 58 video frames in a single input. For long recordings or videos you will chunk content, which is standard practice anyway.
Why Matryoshka truncation matters for devices
The first dimensions of an MRL-trained vector carry most of the information. That lets you keep 768 numbers where quality matters and slice to 256 or 128 where storage and speed matter. Google's guide says 256 dimensions retains roughly 95 percent of multimodal retrieval performance, and the announcement cites up to 6x storage reduction from truncation.
On a phone, an index of a user's photos, notes and voice memos is dominated by vector storage. Cutting each vector to a third of its size is the difference between a feature that fits and one that does not. It also speeds up nearest-neighbor search, which matters when the whole pipeline runs without a server.
Task prompts: queries and documents are embedded differently
The developer guide stresses task-specific prompts. Queries and documents use different prompt prefixes (SearchQuery versus Document) so the model shapes the vector for retrieval rather than generic similarity. This is a common source of bad results with embedding models: people embed the query and the corpus identically and wonder why recall is mediocre. If you adopt EmbeddingGemma 2, follow the prompt format in the model card exactly, and use the same prompt convention at index time and query time.
What the benchmarks say, and what they do not
Google highlights a 78.68 score on MTEB Code (up from 68.76 for EmbeddingGemma 1, roughly a 14 percent relative gain per the developer guide) and "leading scores among sub-1B multimodal embedders" across its benchmarks, including the audio benchmark MAEB.
Read those with the usual care:
- They are vendor-run. Independent leaderboard entries usually appear days or weeks after a release. Watch the MTEB and MAEB leaderboards before you treat the rankings as settled.
- "Sub-1B" is a narrow class. Large API embedders such as Google's own Gemini Embedding family, or Cohere's Embed 5 Pro and Fast, are in different size and cost classes. Our open versus closed embedding model roundup frames how to choose between them.
- Benchmarks are not your data. A model that wins on public code retrieval can still stumble on your internal jargon, your language mix or your noisy audio. Build a 50 to 100 query evaluation set from your real content first.
The code gain is the most useful claim for developers. Google's guide says EmbeddingGemma 2 "significantly outperforms EmbeddingGemma 1 on code search and technical retrieval, making it ideal for local codebase indexing and agentic code search." That targets a real workflow: coding agents and editors that index a repository locally so source code never leaves the machine.
How it fits the Gemma 4 family
EmbeddingGemma 2 is a retrieval companion to the Gemma 4 generation of open models. If you are already running Gemma 4 locally, see our posts on Gemma 4 running in about 2GB of RAM on Apple Silicon and the Gemma 4 updates for tool calling and attention speed. Pair a small generative model with a small embedder and you have the two halves of a fully local RAG system: one finds the relevant material, the other writes the answer. Neither needs a network call, which is the privacy and latency argument Google is leaning on.
There is also a smaller-is-better lineage worth knowing about. Earlier this year developers shipped tiny browser-side embedders, such as the one in our 7MB WebAssembly embedding model guide. EmbeddingGemma 2 is bigger than those, but it adds modalities and Google-grade tooling. Support for transformers.js and WebGPU means it can run in a browser tab as well.
How to try it
The distribution list in Google's announcement is unusually broad, which lowers the cost of a trial:
- Python: load the weights with
transformersorsentence-transformersfrom Hugging Face. - Serving: vLLM and SGLang for batch embedding on a GPU server.
- Mac and local desktops: MLX, Ollama, LM Studio and llama.cpp.
- Mobile and edge: LiteRT, MediaPipe and Google AI Edge.
- Browser: transformers.js with WebGPU.
A sensible first experiment takes an afternoon:
- Pick the smallest configuration that covers your modalities. If you only search documents and code, start at 270M.
- Embed a few thousand of your own items with the document prompt, and 50 realistic queries with the query prompt.
- Compare recall against whatever you use now. Check 768, 256 and 128 dimensions to see where quality drops for your data.
- Measure memory and latency on your real target device, not a laptop.
- Only then plan a full re-index. Vectors from different models are not interchangeable, so migrating means re-embedding everything.
For a RAG pipeline specifically, embeddings are only one piece. Retrieval quality also depends on chunking and on whether you use plain vector search or a reasoning-based approach; our comparison of RAG versus agentic RAG with PageIndex covers the trade-offs.
Why it matters
Three things stand out.
First, one model instead of three. Multimodal products today often glue together separate encoders and reconcile them with custom code. A shared space with a permissive license removes a lot of that plumbing for small teams.
Second, permissive and portable. Apache 2.0 means you can ship it inside a product without negotiating terms, and the modular sizes mean you can fit different devices. For apps in regulated settings, local embedding keeps user photos, recordings and code on the device.
Third, it raises the floor for open embeddings. Closed APIs have led on multimodal retrieval. An open sub-1B model that is competitive on code and audio gives builders a credible alternative, and it pushes API vendors on price. Whether Google's "leading among sub-1B" claim holds up depends on independent testing in the coming weeks.
Caveats before you commit
- Self-reported results. Benchmarks and RAM figures come from Google. Verify them.
- Re-embedding cost. Switching models is a migration, not a config change.
- Context limits by modality. 29 images or 58 frames per input means chunking for video-heavy corpora.
- Cross-modal quality varies by task. Text-to-image retrieval can be strong while audio-to-video retrieval may lag. Evaluate the exact pairs your product needs.
- Safety and privacy still apply. Embeddings can leak information about the content they encode. Treat stored vectors as sensitive data when the source is sensitive.
Bottom line
EmbeddingGemma 2 is a useful, practical release: small enough for a phone, open enough for commercial use, and broad enough to cover the content types most apps actually search. Treat Google's numbers as a starting point, run your own evaluation, and use the modular sizes to load only what you need. If your search or RAG stack today is text-only, this is a reason to look again at what users could find by describing a picture, a sound or a line of code.
Related reading
- What is an embedding? Examples and how it works
- Top 10 open and closed-source embedding models
- Cohere Embed 5 Pro and Fast
- A 7MB browser embedding model with WebAssembly
- Gemma 4 in 2GB of RAM on Apple Silicon
- Gemma 4 updates: tool calling and attention speedups
- RAG vs agentic RAG with PageIndex
- Official: Google's EmbeddingGemma 2 announcement and the developer guide
