EMPVResearch NotesIT
EMPV / RESEARCH NOTESNOTE / 009

TECH NOTE / LOCAL AI

EmbeddingGemma 2 brings multimodal retrieval on-device

Google DeepMind maps text, code, images, video and audio into a shared 768-dimensional vector space. The 740M model is modular and designed for consumer hardware; corpus quality, memory and index cost still need to be tested.

Published2026-10-07 · 7 min
EmbeddingGemma 2Multimodal retrievalRAGLocal AI
01 / The model

One embedding model for text, code, images, video and audio.

Google DeepMind released EmbeddingGemma 2 on October 6, 2026. The official model card describes a 740-million-parameter model that maps text and code, images, video and audio — including combinations in the same input — into a shared 768-dimensional vector space. The weights are available on Hugging Face.

It is not a generative model. It produces comparable numerical representations for semantic search, retrieval-augmented generation (RAG), classification, clustering and similarity. The main change from the first EmbeddingGemma is the move from text embeddings to a natively multimodal index.

02 / One index

A text query can retrieve content that is not text.

Because the modalities share one vector space, a text query can be compared directly with an image embedding, an audio segment or video frames. Google lists an 8,192-token context window and, under the launch configuration, up to roughly 5.5 minutes of audio, 29 images or 58 video frames when using a single modality.

For an enterprise knowledge archive, the pattern is concrete: manuals, visual PDFs, screenshots, technical photographs, recordings, videos and text documentation can share the same retrieval layer. This does not remove segmentation, metadata, permissions or access control, but it can reduce the need to maintain a completely separate representation pipeline for every format.

03 / The footprint

Unused modalities can be left out of the deployment.

The model card separates a 270M text-only component, a 170M vision encoder and a 300M audio encoder. Text plus images uses 440M parameters, text plus audio 570M, and the complete configuration 740M.

In the launch post, Google reports that with quantization on a Pixel 11 Pro the text-only weights require roughly 191 MB of active RAM and the complete multimodal model roughly 567 MB. These are vendor-reported measurements rather than independent benchmarks, but they show the intended target: retrieval on laptops, phones and edge devices rather than exclusively through a cloud API.

This connects directly to the choice between local and cloud AI: generating embeddings close to the corpus can reduce external dependencies and data transfers, but files, vector databases, permissions and indexing pipelines still need their own security boundary.

04 / Index size

From 768 to 256 dimensions: less storage with a measurable trade-off.

EmbeddingGemma 2 uses Matryoshka Representation Learning. Its native 768-dimensional vectors can be truncated to 512, 256 or 128 dimensions and then re-normalized. Moving from 768 to 256 dimensions cuts storage per vector to one third; 128 dimensions gives a 6x reduction.

In Google's published results, multilingual MTEB moves from 61.36 to 60.41 at 256d, MTEB Code from 78.68 to 76.18 and overall MMEB v2 from 59.01 to 56.24. At 128d, MMEB v2 drops to 45.65, and the documentation itself recommends careful validation for multimodal workloads.

There is also a less visible runtime failure mode: the technical documentation recommends `bfloat16` or `float32` and advises against `float16`, because the model can return NaNs or silently degraded embeddings without necessarily raising an explicit error. Numerical precision and vector dimensionality are therefore part of the evaluation, not incidental runtime settings.

05 / Benchmarks

The text improvement is small. The code result moves much more.

The evaluation published by Google places EmbeddingGemma 2 at 61.36 on multilingual MTEB versus 61.15 for the first EmbeddingGemma. On text, the reported improvement is minimal.

On MTEB Code, the score rises from 68.76 to 78.68. If that result transfers to real repositories, the model could be useful for semantic code search, retrieval for coding agents and local codebase navigation. There is no direct EmbeddingGemma 1 comparison for images, visual documents, video or audio because the previous model did not support those modalities.

At launch, most detailed evaluation numbers still come from Google. The model card says the model supports more than 100 languages while also warning that quality may not be uniform across them. For an Italian or domain-specific corpus, public benchmarks do not replace workload-specific evaluation.

06 / The model behind the index

Embedding lock-in stays in the database after the API request is over.

The model card lists an Apache 2.0 license and also states that deployments must comply with the Gemma Prohibited Use Policy. Weight availability has a specific consequence for embeddings: vectors generated today may stay in a database for years, and switching to an incompatible model will normally require re-embedding the corpus.

Simon Willison highlighted this when commenting on the release: even teams that prefer a hosted service may want a model whose weights remain available, so the same model can be run elsewhere if a provider stops serving it.

For an enterprise system, the embedding model should therefore be treated as part of the index schema alongside vector dimensionality, chunking strategy and metadata. Replacing it can look more like a data migration than switching a stateless API.

07 / The enterprise test

The useful question is whether it retrieves the right items from the real corpus.

EmbeddingGemma 2 makes compact local multimodal retrieval plausible, but the evaluation should start with a representative corpus: real documents, PDFs with actual layouts, images, screenshots, recordings, video and code depending on the system being indexed. The test also needs queries for which the expected relevant items are already known.

That makes it possible to compare text-only and multimodal configurations, 768 versus 256 dimensions, recall, latency, memory use and index size. It is the same principle from our note on evaluating a local LLM before production: model, runtime and workload need to be measured together.

The advantage of local execution is not simply avoiding the cloud. It is the ability to place model, corpus and index inside a defined operating boundary and measure, with reproducible tests, what retrieval quality the whole system can sustain over time.

EMPV / TAKEAWAY

EmbeddingGemma 2 can place text, code and media inside one vector space on local hardware. The useful decision starts with the corpus, real queries and the long-term cost of maintaining the index.

EMPV / SHARE

Share this Research Note.

Research Notes