Google DeepMind Launches EmbeddingGemma 2 Multimodal On-Device Embedding Model
Alphabet's (GOOGL) Google DeepMind announced EmbeddingGemma 2, an open multimodal embedding model that expands the original text-focused EmbeddingGemma to code, images, video and audio, mapping all data types into a shared 768-dimension vector space for on-device cross-modal search and retrieval. Built on the Gemma 4 architecture with 740 million parameters under an Apache 2.0 commercial license, the model is designed to run locally on phones and PCs rather than in the cloud. Its modular design uses 270 million parameters for text and code, 440 million with the visual encoder, 570 million with the audio encoder, and the full 740 million for all modalities. Context length quadruples to 8,192 tokens. Google said the MTEB Code score rose to 78.68 from 68.76, a roughly 14% gain. Matryoshka Representation Learning lets developers shrink vectors to 512, 256 or 128 dimensions, cutting storage needs by up to about sixfold; quantized text mode needs about 191MB of active RAM on Pixel 11 Pro, versus 567MB for the full multimodal version. Weights are available on Hugging Face and Kaggle.
General financial information, not personalized investment advice. Financial disclaimer