Google DeepMind releases EmbeddingGemma 2 for local search across text, images, audio and video
The 740-million-parameter model adds media inputs to Google's earlier text embedding model. Its weights are available under Apache 2.0, while performance and device memory figures remain company claims.
Google DeepMind announced EmbeddingGemma 2 on October 6, 2026, releasing a model designed to place text, code, images, audio and video in one searchable representation. Its weights are available to developers under an Apache 2.0 license. The expansion matters for applications that need to find related material across different kinds of files, including on consumer devices, although Google has not established how those applications will perform across users' hardware.
The release extends EmbeddingGemma, which handled text. Google describes the new model as a way to find a video clip from a voice memo or search audio recordings using a text query. Those are examples of what developers could build, rather than evidence that a finished consumer application has delivered those results.
How EmbeddingGemma 2 handles different media
An embedding turns an input into numbers that software can compare for similarity. According to Google's model card, EmbeddingGemma 2 maps inputs and combinations of inputs into a shared space of 768 dimensions. A search system could use those representations to compare a written query with a relevant image, recording or video segment, instead of relying only on matching words in a file name or transcript.
Google says the model has 740 million parameters and is built on its Gemma 4 architecture. The model card breaks that total into a 270-million-parameter text model, a 170-million-parameter vision encoder and a 300-million-parameter audio encoder. Developers can load the components needed for their application, so a text-only use does not require loading the full set of media components.
The model card lists an 8,192-token context window. It also says developers can reduce the 768-dimensional output to 512, 256 or 128 dimensions to cut the storage needed for embeddings. Google's guidance describes minimal quality impact down to 256 dimensions and identifies 128 dimensions as suitable for text-only workloads. Those options give developers choices about storage and model configuration; their effect on a particular search task still depends on that task.
What developers can download now
Google's launch post says the model weights are available on Hugging Face and Kaggle. It describes availability through Gemini Enterprise Agent Platform Model Garden as coming soon, without establishing that it is available there now. Hugging Face's Transformers documentation says EmbeddingGemma 2 was added on October 6 and describes its shared 768-dimensional space for cross-modal inputs.
The Google model card identifies the license as Apache 2.0. The Expectancy separately reported checking the Hugging Face model listing and license, and counted 744,371,512 checkpoint parameters, consistent with Google's rounded 740-million figure. The license and downloadable weights are concrete parts of this release. They do not, by themselves, establish how accurately an application will retrieve a user's files.
Google's benchmark and device claims
Google reports that EmbeddingGemma 2 scored 78.68 on MTEB Code, up from 68.76 for EmbeddingGemma 1. Its reported multilingual MTEB score is 61.36, compared with 61.15 for the earlier model. The model card also lists results for image, video and audio retrieval, including 64.64 on MIEB lite, 50.67 Hit@1 on MMEB v2 Video and 69.54 MRR@10 on MSEB Retrieval. These are Google's reported benchmark results; the reviewed reporting does not establish an independent reproduction of them.
For a quantized model on a Pixel 11 Pro, Google reports about 191 MB of active RAM for text-only use and about 567 MB for the full multimodal configuration. Those measurements apply to the company's stated device and setup. They should not be read as a memory or performance guarantee for other phones and laptops. The Expectancy likewise cautioned that it had not independently verified Google's performance or memory figures.
From text embeddings to on-device media search
The previous EmbeddingGemma was a lightweight text embedding model based on Gemma 3, according to its 2025 research paper. Google's new release moves the product line from text to a wider range of media. Google says the predecessor recorded more than 20 million downloads, a company figure about the earlier model rather than evidence of adoption for EmbeddingGemma 2.
Google presents local search and retrieval-augmented generation as intended uses. Running a suitable application on a device could keep processing local, but the release materials do not establish that every application built with the model will work entirely offline or keep all user data on the device. Developers and users will need to judge those properties in the application itself. For now, the confirmed change is the model release and its documented capabilities; broad device performance, retrieval accuracy on personal collections and uptake remain open questions.
Sources and context
- EmbeddingGemma 2: an open, lightweight multimodal embedding modelGoogle DeepMind
- EmbeddingGemma 2 model cardGoogle AI for Developers / Google DeepMind
- EmbeddingGemma2Hugging Face Transformers documentation
- Google’s new open embedding model can search your voice memos and videos on-device, without uploading anythingThe Expectancy
- EmbeddingGemma: Powerful and Lightweight Text RepresentationsarXiv (Google-authored research paper)
AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.
About NewsJaws Desk
AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.