EmbeddingGemma 2 is Google DeepMind's new open embedding model, released under Apache 2.0, that turns text, code, images, video and audio into vectors in one shared space. For teams building retrieval for AI agents, it makes private, on-device and multimodal search cheaper to try, but moving an existing index means re-embedding everything.
What was released
EmbeddingGemma 2 was announced by Google DeepMind on 6 October 2026 in the post EmbeddingGemma 2: an open, lightweight multimodal embedding model. The weights are available today on Hugging Face and Kaggle under the Apache 2.0 licence, and Google says Vertex AI Model Garden support is coming soon. Use is also subject to the Gemma Prohibited Use Policy, as the model card notes.
It is built on Gemma 4 and works with the tools most teams already use for retrieval: sentence-transformers, transformers, vLLM, SGLang, llama.cpp, Ollama, LM Studio and MLX, plus LiteRT, MediaPipe and transformers.js for on-device and in-browser use.
What actually changed
- Multimodal input. The first EmbeddingGemma handled text only. Version 2 embeds text, code, images, video and audio, and the developer guide says all of them land in the same compatible vector space, so a text query can retrieve an image or an audio clip directly.
- Modular encoders. The text and code core is small, and the vision and audio encoders are optional. Google gives the sizes as 270M parameters for text only and 740M with every encoder loaded. The developer guide says disabled encoders are never loaded into memory, and that a text-only index can later take image or audio embeddings without recomputing the text vectors.
- A longer input window. The model card lists an 8,192-token context, which Google says is larger than the first version's. Per input, that also covers roughly 29 images, 58 video frames or about 327 seconds of 16 kHz mono audio at the default settings.
- Better code retrieval. Google reports a gain on the MTEB Code benchmark from 68.76 to 78.68 over the first version. That is Google's own result on a public benchmark, not a measure of your codebase.
- Smaller vectors when you need them. Outputs are 768 dimensions and can be cut to 512, 256 or 128 with Matryoshka Representation Learning, so the same model can serve a high-recall index and a compact one. It covers more than 100 languages.
- Runs on a phone. Quantised on a Pixel 11 Pro, Google measured about 191MB of active RAM for the text weights and 567MB for the full model.
What it means for teams building AI products and agents
In CodeDTX's view, the release matters most where data cannot leave a device or your own servers, and where knowledge lives in more than text.
- Retrieval that stays private. An Apache 2.0 model this small can embed documents on a laptop, a phone or a single server, so sensitive content never goes to a hosted embedding API. That fits teams with the residency constraints covered in data residency for AI agents, and pairs with a self-hosted generator, as in running an AI agent on a model you host yourself.
- One index for screenshots, recordings and documents. Support knowledge often sits in screen captures, call recordings and diagrams. A shared vector space lets an agent search all of it with one query and one index, instead of transcribing or captioning everything first. Load only the encoders you need, because each one adds memory.
- Cheaper storage at scale. Truncating to 256 or 128 dimensions shrinks the vector store and speeds up search. Measure recall on your own queries at each size before you commit, and keep the truncation choice in your index metadata so query and document vectors always match.
- Get the prompts right. The model expects task prefixes on text, such as separate query and document formats for retrieval; images, video and audio need none. The model card also says to run it in bfloat16 or float32, not float16. Both are easy to miss in a quick port and quietly lower retrieval quality.
When not to switch
Moving from EmbeddingGemma 1 or a hosted embedding API is a migration, not a model swap. Vectors from different models are not comparable, so you must re-embed the whole corpus, run old and new indexes side by side, and switch only when your own retrieval evals show a gain; see what AI agent evals catch. If your agent searches only text and the current index already returns the right passages, the cost of re-indexing may not pay back. Teams that rely on Vertex AI should wait for the managed endpoint rather than run the model themselves in the meantime. And keep the old index until the new one has served real traffic, as in rolling out an AI agent in stages.
Frequently asked questions
What is EmbeddingGemma 2?
EmbeddingGemma 2 is an open embedding model from Google DeepMind, announced on 6 October 2026. It converts text, code, images, video and audio into vectors in one shared space, so they can be searched together. It is built on Gemma 4, released under the Apache 2.0 licence, and small enough to run on a phone or laptop as well as a server.
Where can I download EmbeddingGemma 2?
The weights are on Hugging Face as google/embeddinggemma-2 and on Kaggle. Google says support in Vertex AI Model Garden is coming soon. You can load it with sentence-transformers or transformers, serve it with vLLM, SGLang, llama.cpp, Ollama or LM Studio, and run it on device with LiteRT, MediaPipe or transformers.js. Check the model card for the required task prompts first.
Do I need to re-embed my documents to use EmbeddingGemma 2?
Yes, if you are moving from EmbeddingGemma 1 or any other embedding model. Vectors from different models live in different spaces, so queries embedded with the new model will not match documents embedded with the old one. Build a new index alongside the current one, compare retrieval quality on real queries, then switch. Within version 2, adding image or audio encoders later does not require recomputing text vectors.
Can EmbeddingGemma 2 run on device?
Yes. Google designed it for on-device use and reports that the quantised text-only weights use roughly 191MB of active RAM on a Pixel 11 Pro, with the full multimodal model needing more. It runs with LiteRT, MediaPipe and transformers.js, including in the browser with WebGPU. That makes private search over local files possible without sending content to a hosted API.
Is EmbeddingGemma 2 good for code search?
Google reports a clear gain on the MTEB Code benchmark over the first version, which makes it worth testing for coding agents and internal code search. Benchmarks are not your repository, though. Build a small set of real questions with known answers from your own code, compare it against your current embedding model, and check results at the vector size you plan to store.



