Google releases a 740M-parameter embedding model for local multimodal search
EmbeddingGemma 2 maps text, code, images, video and audio into one vector space; Google says its weights are available under Apache 2.0.
By Ryan Merket · Published
Primary source: X
Why it matters
Google is offering developers a downloadable, Apache 2.0 model for cross-media retrieval that can run locally, alongside its cloud embedding API. Developers must weigh device memory and retrieval quality against keeping data on-device.

Google released EmbeddingGemma 2, an open-weight model designed to let developers search across text, code, images, video and audio on local devices. In its October 6th X thread, Google described it as a multimodal embedding model built on Gemma 4 and licensed under Apache 2.0. The model card and developer guide add the operating detail: different kinds of input are mapped into the same 768-dimensional space, so a text query can be compared with an image, video frame or audio clip.
https://x.com/Google/status/2107505117511528640
Retrieval often depends on a chain of separate models. Google's AI Edge post says EmbeddingGemma 2 can replace parts of that chain, such as image captioning, speech-to-text and text embedding, with local vector search. Its examples include searching a phone's videos for a particular moment without transcribing the audio or generating captions, and finding an image with a natural-language query. Google says those searches run on-device, without an internet connection.
The 740-million-parameter total is modular, rather than a requirement to load every component. Google's guide describes a 270-million-parameter text-and-code configuration, a 440-million-parameter text-and-vision version, a 570-million-parameter text-and-audio version, and the full 740-million-parameter model. The architecture pairs those components in one embedding space, allowing a text-only query to be matched against items embedded with the larger multimodal setup. Google reports active RAM of about 191MB for text-only use and 567MB for the full model on a Pixel 11 Pro; those are Google's measurements, not independently verified device tests.
The shared space also has limits developers will need to account for. The model card sets a shared 8,192-token context budget and estimates that, when using only one modality, it can cover about 5.5 minutes of audio, 29 images or 58 video frames. Mixing media consumes that same budget, and video is sampled at one frame per second by default. Google also supports shortening embeddings from 768 dimensions to 512, 256 or 128 to save storage. Its own evaluation shows the trade-off: at 128 dimensions, the multimodal benchmark score falls from 59.01 to 45.65, while the code benchmark falls from 78.68 to 71.41. The company recommends 128 dimensions mainly for text-only workloads.
Google's benchmark results make the case most clearly for code retrieval. The model card reports a score of 78.68 on MTEB Code, compared with 68.76 for the first EmbeddingGemma. On the multilingual MTEB benchmark, the listed scores are 61.36 and 61.15, respectively, a much smaller difference than the code-retrieval gap. These are figures published by Google; the company describes the benchmarks and evaluation setup in its model card. The results give developers a basis for comparison, but they do not establish how the model performs on any particular private collection of photos, recordings, documents or code.
The weights are available through Hugging Face, and Google also lists Kaggle as a distribution option. The company has published integrations and examples using Sentence Transformers, Transformers and its own MediaPipe and LiteRT tooling. Its AI Edge Gallery includes demos for searching local media and locating moments in video. Google says support through Android's ML Kit is planned for the coming weeks, placing that managed integration after today's downloadable release.
The product gives Google a local counterpart to Gemini Embedding 2, which the company documents as a multimodal model offered through the Gemini API. EmbeddingGemma 2 puts the weights in developers' hands for local deployment; Google's cloud API offers a different route for embedding work. Developers handling personal media or data that should remain on a device can use local deployment, while weighing local hardware and indexing work against a hosted service. Adoption will depend on whether the smaller, modular model and Google's edge software support cross-media search beyond demonstrations.