AI · Oct 6, 2026
EmbeddingGemma 2: an open, lightweight multimodal embedding modelGoogle DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4
Oct 6, 2026, 11:36 AM · MarkTechPost

EmbeddingGemma 2 brings search across photos, video, audio and code onto a phone under a permissive license, which is great for privacy and a little unsettling for the same reason.
Why it matters
Google DeepMind has released EmbeddingGemma 2, an open model that turns text, code, images, video and audio into points in one shared 768-dimensional space, so a typed question can pull up a photo and a voice memo can find a video clip. It has 740 million parameters, an 8,192-token context window that MarkTechPost describes as four times the first version, and an Apache 2.0 license. Weights are on Hugging Face and Kaggle, with builds for Ollama, llama.cpp and Google's on-device runtime.
The practical shift is size. Google reports about 191 megabytes of active memory for the quantized text-only setup on a Pixel 11 Pro and about 567 megabytes for the full multimodal model. That puts private, offline search over a person's own media within reach of an ordinary phone app.
From the desk
We think this is a genuinely useful release, and the license is the underrated part. The first EmbeddingGemma shipped under Gemma's own terms; this one is Apache 2.0, which is about as simple as commercial use gets. Pair that with a modular design, where developers load a 270 million parameter text core and add vision or audio encoders only if they need them, and you get a model builders can actually ship. Every setup shares one vector space, so a query embedded with the small text-only version can match documents indexed by the full model. That is a smart piece of engineering for apps that run on different devices.
On quality, the gains are concentrated where developers will notice. Google's model card, as reported, shows code retrieval rising from 68.76 to 78.68 on MTEB Code, while multilingual text barely moves, from 61.15 to 61.36. These are Google's own numbers. MarkTechPost also notes that larger models such as Qwen3-VL-Embedding-2B report a higher score on one multimodal benchmark, with roughly 2.7 times the parameters and no audio. So this is a strong small model, not the best embedder at any size.
Now the other side. On-device retrieval is the privacy-first pattern we have argued for: your data stays on your hardware and does not need a cloud call. But making every photo, clip and recording on a phone instantly searchable by meaning also makes it easier for any app with access to sift a person's media for things they never meant to surface. The same capability that finds your vacation video finds a screenshot you forgot you took. If this pattern scales, operating system permissions for media access will matter far more than they do today.
There is a strategic wrinkle too. Google sells Gemini Embedding 2 as a paid API, and here it is giving away a capable small alternative. We read that as Google betting on the on-device ecosystem, with Android ML Kit support said to be coming within weeks, rather than leaving that space to rivals. Good for developers. Worth watching for how the two products end up positioned.
Context
Embedding models convert content into lists of numbers that capture meaning, so similar items sit close together and can be found quickly. They are the retrieval half of most retrieval-augmented generation systems. EmbeddingGemma 2 also supports shortening its vectors to 512, 256 or 128 dimensions to save storage; Google's figures show little loss at 256 for multilingual text but a sharp drop on multimodal tasks at 128, so it recommends the smallest size mainly for text-only use.
Who feels it
- App developers
- A permissively licensed multimodal embedder small enough for phones lowers the bar for offline search, tagging and private RAG features.
- Phone users
- Better local search over personal media is likely. So is the need to think harder about which apps get access to photos and recordings.
- Enterprises with sensitive data
- Local embedding keeps documents and media off third-party servers, a meaningful option for regulated or privacy-bound teams.
- Competing model makers
- An Apache-licensed small model with audio support raises the baseline other open embedders will be measured against.
What to watch
- Independent benchmark runs that confirm or complicate Google's self-reported scores
- Arrival of Android ML Kit support with NPU acceleration, which Google says is weeks away
- How phone platforms handle permissions as apps adopt on-device semantic search of media
- Whether Google adjusts Gemini Embedding 2 pricing or positioning in response
Companies: Google