I love it when a release quietly makes a whole category of apps possible, and EmbeddingGemma 2 feels like one of those. Google released it on October 6 as an open embedding model that understands text, code, images, video and audio, small enough to run on a phone. If you’ve ever wanted to search your own photos, voice notes or codebase without anything leaving your device, let’s walk through what it does and how to start building with it.

What an embedder does
An embedding model turns a piece of content into a list of numbers, called a vector. Content with similar meaning lands close together in that space. “A dog running on the beach” and a photo of a dog running on the beach end up as near neighbours, so you can find one by searching with the other.
That’s the heart of search in retrieval augmented generation (RAG). You embed your notes, files or photos once and store the vectors. When a question comes in, you embed the question too, then pull back the closest matches. Those matches become the context for a small language model to answer from, or the candidates for a decision model to choose between, which I covered in my post on decision models. The embedder finds, and the model on top picks or answers.
When every step runs locally, your data stays on your device. That’s why a good small embedder matters so much for private, on-device AI.
What’s new in EmbeddingGemma 2
According to Google’s launch post, EmbeddingGemma 2 is built on Gemma 4 and maps text, code, images, video and audio into one shared space. The developer guide puts that space at 768 dimensions.
The part I like most is how modular it is. You load only what you need:
- 270M parameters for text and code
- 440M with the vision encoder, for images, visual documents and video frames
- 570M with the audio encoder
- 740M for everything
All of these setups share the same vector space. Google’s guide says a query embedded with the text-only setup “can be matched directly against documents embedded with the full model,” and adding an encoder later doesn’t mean re-computing what you’ve already stored.
A few more details worth knowing:
- License: Apache 2.0, a commercially friendly change from the Gemma license on the first EmbeddingGemma. The model card still asks deployments to follow the Gemma Prohibited Use Policy.
- Context: 8K tokens, which Google says is four times the first version. That’s enough for up to 5.5 minutes of audio, 29 images or 58 video frames.
- Matryoshka truncation: you can cut vectors from 768 dimensions to 512, 256 or 128. Google says that’s up to a 6x storage saving, and the guide gives a nice example: a million 768-dimension vectors take roughly 1.5 GB in bfloat16, and about 250 MB at 128.
- Shared parts with Gemma 4: the two models share a text tokenizer and audio encoder, so running them together in one pipeline costs less memory.
The numbers, and whose they are
Every quality and memory figure here comes from Google. With quantization on a Pixel 11 Pro, Google reports about 191 MB of active RAM for text only and about 567 MB for the full multimodal model.
On quality, the standout is code. Google reports MTEB Code rising from 68.76 to 78.68. The model card shows multilingual text holding steady at 61.36, against 61.15 for the first version. So if you’re already happy with EmbeddingGemma for text search, expect similar text results, with code retrieval, new modalities and an open license as the real gains.
The card’s truncation table is useful for planning. At 256 dimensions, MTEB multilingual drops only slightly, from 61.36 to 60.41. At 128 dimensions, the overall multimodal score falls from 59.01 to 45.65. That matches Google’s own advice to save 128 dimensions mostly for text.
I haven’t seen an independent MTEB entry for EmbeddingGemma 2 yet, and Google hasn’t published latency or battery figures.
It shipped everywhere on day one
This is the part that makes me want to start building right away:
- LiteRT-LM v0.18.0 adds EmbeddingGemma 2 across Python, Kotlin, Swift, Web and C++, plus an OpenAI-compatible
/v1/embeddingsendpoint inlitert-lm serve. - transformers.js 4.3.1 runs it in the browser on WebGPU or WASM, using
onnx-community/embeddinggemma-2-ONNX. - llama.cpp users can grab GGUF builds from unsloth and ggml-org.
- Google AI Edge Gallery has on-device demos, including Instant Media Search and Video Moments Finder.
- Ollama has an embeddinggemma-2 library page with 270m, 440m, 570m and 740m tags.
On the Mac side, Ollama v0.40.0 is now the latest release. On Apple Silicon, supported models run on MLX by default, including qwen3.8, gemma4, qwen3.6, qwen3.5 and the Nimble, Tev1, Clef and Clef-Flash decision models. The notes add, “MLX now has support for an embedding model: embeddinggemma-2.” Ollama hasn’t published speedup numbers, so this is a great one to measure yourself.
Getting a first vector takes a minute. This comes straight from Ollama’s library page:
ollama pull embeddinggemma-2
curl http://localhost:11434/api/embed \
-d '{
"model": "embeddinggemma-2",
"input": "Why is the sky blue?"
}'
Try it yourself
Here’s how I’d get started:
- Pick the module that matches your data. Text and code need only the 270M core. Add vision for photos and screenshots, and audio for voice notes or recordings.
- Experiment with Matryoshka dimensions. Try 256 dimensions for a storage-friendly index, and check recall against 768 on your own content.
- Build a small local index. A few hundred of your own photos, notes or source files is plenty to get a feel for cross-modal search.
- Measure on your hardware. Time your embeddings and watch battery use on your phone or laptop, because those are the numbers that matter for your app.
- Watch for independent results. Once EmbeddingGemma 2 shows up on public MTEB, compare it against your own tests.
Search that runs privately on your own device, across everything you’ve captured, is a lovely thing to be able to build. Pick a folder, embed it, and see what you find.