One embedding model for text, code, images, video and audio.
Google DeepMind released EmbeddingGemma 2 on October 6, 2026. The official model card describes a 740-million-parameter model that maps text and code, images, video and audio — including combinations in the same input — into a shared 768-dimensional vector space. The weights are available on Hugging Face.
It is not a generative model. It produces comparable numerical representations for semantic search, retrieval-augmented generation (RAG), classification, clustering and similarity. The main change from the first EmbeddingGemma is the move from text embeddings to a natively multimodal index.