Google rolls out EmbeddingGemma 2 with multimodal on-device embeddings


EmbeddingGemma 2

Google has launched EmbeddingGemma 2, an open and lightweight multimodal embedding model that maps combinations of text, images, audio and video into a unified embedding space. The model expands EmbeddingGemma beyond text to support code, images, video and audio.
Google introduced EmbeddingGemma last year as a lightweight option for high-quality text embeddings, helping apps organize, search and connect information directly on consumer hardware. The model has recorded more than 20 million downloads, with developers using it for on-device search tools and privacy-focused retrieval augmented generation (RAG) pipelines.

EmbeddingGemma 2

EmbeddingGemma 2 is built on the Gemma 4 architecture and has 740 million parameters. It is released under the Apache 2.0 license and can process different types of content through a single natively multimodal model. For example, it can find a specific video clip from a voice memo or search hours of audio recordings using a text query.

Built using the same technology as Gemini Embedding models, EmbeddingGemma 2 includes the following capabilities:

  • Text-only workloads: Requires as little as 270 million parameters.
  • Vision: Includes an optional 170-million-parameter vision encoder.
  • Audio: Includes an optional 300-million-parameter audio encoder for full multimodal support.
  • Storage and memory: Uses Matryoshka Representation Learning (MRL) to dynamically truncate output vectors from 768 dimensions to 512, 256 or 128 dimensions. This provides up to 6x storage reduction for local vector databases and memory usage.
  • On-device inference: With quantization, it requires as little as around 191MB of active RAM for text-only weights and around 567MB for the full multimodal model on a Google Pixel 11 Pro.
  • Context window: Provides an 8K-token context window, which is four times larger than EmbeddingGemma 1. It can process up to 5.5 minutes of audio, 29 images, 58 video frames or interleaved combinations of these inputs directly on local hardware.

Code, vision and audio performance

EmbeddingGemma 2 matches the multilingual text performance of EmbeddingGemma while improving code performance. Its MTEB Code score increases by 9.92 points, from 68.76 to 78.68.

Google says EmbeddingGemma 2 achieves leading scores among sub-1B multimodal embedding models for its size across benchmarks including MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark). It also matches or outperforms many larger models across text, vision and audio tasks.

Across image, video, document and audio workloads, Google says the model improves quality per parameter for sub-1B models and outperforms some specialist models that are more than twice its size. The code performance supports local codebase indexing, semantic code search and coding agent retrieval.

EmbeddingGemma 2 Performance

On-device semantic search and retrieval

EmbeddingGemma 2 generates embeddings locally, which helps with data privacy and reduces pipeline latency. It also enables cross-modal search and retrieval that can operate entirely offline.

When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines for complex multimodal data. Since it is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.

Availability and developer tools

EmbeddingGemma 2 is available through the following platforms and tools:

  • Model downloads: Model weights are available on Hugging Face and Kaggle, while Gemini Enterprise Agent Platform Model Garden availability is coming soon. Models optimized for on-device use are available through LiteRT Community on Hugging Face.
  • On-device deployment: Google AI Edge MediaPipe supports embedding, retrieval and decision tasks, while LiteRT supports custom model integration. Developers can also build for the browser using transformers.js or WebGPU.
  • Development and serving: The model can be served using transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio. Qdrant can be used to store embedding vectors.
  • Fine-tuning: Developers can follow guidance from Unsloth to fine-tune EmbeddingGemma 2 for their use cases.