Google is bringing multimodal semantic search a lot closer to the hardware with EmbeddingGemma 2, a new open-weight embedding model designed to understand text, images, video and audio without sending that data to the cloud.
Announced October 6, EmbeddingGemma 2 is the successor to last year’s text-focused EmbeddingGemma. The new model expands the family into native multimodal embeddings while keeping its footprint small enough for phones, laptops and other consumer devices.
The key idea isn’t to have the model generate an answer about your files. Instead, EmbeddingGemma 2 turns different types of content into numerical representations called embeddings. Content with similar meaning ends up close together in a shared vector space, allowing software to search, retrieve, classify and organize information based on meaning rather than just matching keywords.
That becomes particularly interesting when the content isn’t all text.
One embedding space for text, images, video and audio
EmbeddingGemma 2 can map text, code, images, video and audio into the same 768-dimensional embedding space. That means developers can build cross-modal search systems where a query in one format can retrieve information stored in another.
For example, a developer could use a text description to find relevant images, search a video library for a particular scene, or use an audio recording to locate a corresponding moment in a video.
Google gives a particularly useful example: finding a specific video clip using a voice memo. Another example is searching through hours of audio recordings using a text query. Both can be handled by the same natively multimodal model rather than stitching together separate models for speech recognition, image understanding and text embeddings.
The model can also handle interleaved multimodal inputs. Its 8,192-token context can contain combinations of text, images, video and audio, allowing developers to create a single embedding representing a mixture of media.
That could be useful for everything from personal media libraries to product catalogs and document retrieval systems.
It is designed to run locally
The other half of the EmbeddingGemma 2 story is where it runs.
Google has designed the model for on-device inference, rather than assuming developers have access to a large cloud GPU. The full model contains 740 million parameters, but its architecture is modular. The text component has 270 million parameters, while optional vision and audio encoders add 170 million and 300 million parameters, respectively.
That means an application doesn’t necessarily need to load the entire model.
A text-only application can use the smaller configuration, while developers working with images or audio can add the relevant encoder. The full multimodal configuration reaches 740 million parameters.
Google says that, after quantization, EmbeddingGemma 2 can use around 191MB of active RAM for text-only weights and around 567MB for the full multimodal model on a Pixel 11 Pro. Those figures illustrate why Google is positioning it for phones and other resource-constrained hardware.
The local approach also has an important privacy advantage. Embeddings can be generated directly on the device, allowing applications to search personal media or documents without necessarily uploading the underlying content to a remote server. It can also reduce latency and allow retrieval features to work offline.
A bigger context window and smaller vectors
EmbeddingGemma 2 also gets a substantially larger context window than its predecessor.
The model supports an 8K-token context window, which Google says is four times larger than the first EmbeddingGemma. In practical terms, the model can process up to about 5.5 minutes of audio, 29 images, or 58 video frames, along with combinations of those inputs.
Google is also using Matryoshka Representation Learning to give developers more flexibility over the size of the resulting embeddings.
The default output is 768 dimensions, but developers can truncate it to 512, 256 or 128 dimensions. Smaller vectors take less storage and can make similarity searches more efficient, although there is a corresponding trade-off in quality. Google says this approach can reduce local vector database storage and memory requirements by as much as six times.
That’s a meaningful detail for applications that need to index thousands or millions of pieces of content locally. The difference between storing 768-dimensional vectors and much smaller representations can add up quickly.
Google says the model punches above its size
EmbeddingGemma 2 is based on the Gemma 4 architecture and is released under the Apache 2.0 license, making it available for developers to use and adapt under a commercially permissive license.
Google describes it as a best-in-class model for its size, citing performance across text, code, vision and audio benchmarks.
One particularly notable improvement is code retrieval. Google reports that EmbeddingGemma 2 scores 78.68 on MTEB Code, compared with 68.76 for the original EmbeddingGemma, a 9.92-point improvement.
That makes the model relevant beyond photo and media search. Developers could use it for local codebase indexing, semantic code search and retrieval systems used by coding agents.
The model card also documents support for multilingual text and the ability to use shorter embedding dimensions when storage efficiency matters.
EmbeddingGemma 2 could make local search much smarter
The interesting part of EmbeddingGemma 2 isn’t simply that another small AI model has arrived. It’s that embedding technology is increasingly becoming something developers can put directly into consumer applications.
Traditional search often depends heavily on text metadata. If a photo isn’t labeled properly, a keyword search may have no idea what is inside it. A semantic embedding model can instead represent the meaning or characteristics of the content and compare it with the meaning of a query.
Making that process multimodal opens the door to more natural searches.
Imagine searching a phone’s media library by describing what you’re looking for rather than remembering a filename. Or searching recorded meetings using a text description. Or finding a particular portion of a video from a short audio note.
Google is already demonstrating these ideas through its AI Edge tools, including Instant Media Search and Video Moments Finder. EmbeddingGemma 2 can also be paired with Gemma 4, with EmbeddingGemma handling local retrieval and Gemma handling the contextual reasoning once relevant information has been found.
That separation is important. An embedding model doesn’t need to be a general-purpose chatbot. Its job is to make finding the right information fast and efficient, after which a larger generative model can take over if necessary.
Google says EmbeddingGemma 2 is available through Hugging Face and Kaggle, with deployment options including Google AI Edge and LiteRT. Developers can also work with tools such as Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang and Ollama.
For developers building privacy-focused search, retrieval-augmented generation or media applications, that makes EmbeddingGemma 2 an interesting piece of infrastructure. Instead of treating multimodal search as a cloud-only capability, Google is pushing the technology toward the devices where the data actually lives.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
