SDSignal Desk

EmbeddingGemma 2: an open, lightweight multimodal embedding model

Oct 6, 2026, 12:57 PM · Google DeepMind

Image: Google DeepMind

Google's small open embedding model makes private, offline search of photos, voice memos and video practical on a phone, which is a privacy win and a new kind of power.

Why it matters

Google DeepMind released EmbeddingGemma 2, a 740 million-parameter open model under the Apache 2.0 license that maps text, code, images, audio and video into one shared embedding space. Embeddings are what let software find things by meaning rather than exact words, so this is the plumbing behind search, retrieval and recommendation.

What changed is where that plumbing can live. Google says that with quantization the full multimodal model needs about 567MB of active memory on a Pixel 11 Pro, and about 191MB for text alone. It has an 8K-token context window, four times the first version, enough for roughly five and a half minutes of audio or 58 video frames. The first EmbeddingGemma passed 20 million downloads.

From the desk

We like this release a lot, and for a plain reason: it moves useful AI onto hardware people already own. Searching hours of recordings with a text query, or finding a video clip from a voice memo, used to mean shipping your media to someone's server. A model this small, running offline, keeps that data on the device. For developers who care about privacy, the permissive license and modular design matter too. A text-only app needs as little as 270 million parameters, with vision and audio encoders bolted on only when needed.

The practical details back up the pitch. Google reports code-retrieval scores on the MTEB Code benchmark rising from 68.76 to 78.68, which makes local codebase search and coding-agent retrieval more realistic. Vector sizes can be trimmed from 768 dimensions down to as few as 128, cutting storage up to sixfold. These are Google's numbers, and we'd want independent tests across messy real-world media, but the direction is clear.

Here's the part we don't want lost. Cheap, private, on-device search across every photo, recording and video is also cheap, private, on-device search for whoever controls the device. Stalkerware, abusive partners, or anyone who gets hold of a phone could sift a person's media by meaning in seconds, with no server logs to show for it. Privacy from the cloud is not the same as privacy from the person next to you. As this capability becomes standard in apps, we'll be watching whether platforms add guardrails around who can index what.

Our read: a strong, practical open release that pushes AI search toward the edge, where we think it belongs. The responsibility shifts to app makers to use it well.

Context

Google says EmbeddingGemma 2 is built on the Gemma 4 architecture and shares its text tokenizer and audio encoder, so the two can run together on one device with a smaller combined memory footprint. Weights are on Hugging Face and Kaggle, with support across common tools such as transformers, llama.cpp, Ollama and vLLM.

Who feels it

App developers
A small, permissively licensed model for offline multimodal search and retrieval, with tooling support from day one.
Phone and laptop users
Smarter search of personal media without uploading it, if apps adopt it.
Privacy and safety advocates
On-device indexing reduces cloud exposure but makes deep searches of a seized or shared device far easier.

What to watch

  1. Independent benchmarks of multimodal retrieval quality on real personal media
  2. Which consumer apps ship on-device search built on it
  3. Availability in Google's enterprise Model Garden, which the company says is coming soon
  4. Platform-level controls on which apps can index photos, audio and video

Read the original

Continue at the source.

Google DeepMind

Also covering this