Build4 publishers3 min readPublished
Google's EmbeddingGemma 2 fits multimodal search into 567MB of phone RAM
Google released EmbeddingGemma 2, an Apache 2.0 model that maps text, code, images, video and audio into one embedding space with 740M parameters. Phone apps get offline cross-media search from one set of weights, within limits set by a shared 8K-token window and lossy vector truncation.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A text-and-code core of 270M parameters runs on its own, with optional 170M vision and 300M audio encoders added for full multimodal support.
- With quantization on a Pixel 11 Pro, Google measured about 191MB of active RAM for text-only use and about 567MB for the full multimodal model.
- MTEB Code rose to 78.68 from 68.76 on the first EmbeddingGemma, and Google pitches the model for local codebase indexing and coding-agent retrieval.
- Weights are on Hugging Face and Kaggle now, and Google says Android ML Kit support is planned for the coming weeks.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability According to Google's AI Edge post, a phone app can find a photo or a video moment from a text query offline without first running separate captioning and speech-to-text models.
- cost Mixed-media indexes have to keep long vectors, because cutting to 128 dimensions costs cross-modal retrieval about two and a half times the share of quality it costs code.
- constraint Long recordings and videos must be split by the app before embedding, so chunk length and overlap become indexing decisions the developer owns.
- decision Teams holding personal media now choose between local weights and Google's hosted Gemini Embedding 2, trading device memory and indexing work for keeping data on the device.
Every input type lands in the same 768-dimensional space, so a text query is scored directly against an image, a video frame or an audio clip [8]. The encoders are separable at query time. RuntimeWire, citing Google's guide, reports that a text-only query can be matched against items embedded with the full multimodal setup [5]. In my view that split is the best decision in the release. A search box can run on the 270M text core, and the vision and audio encoders, about 376MB of extra RAM on Google's Pixel 11 Pro figures, need loading only when new media is being indexed [5][23]. RuntimeWire notes those memory figures are Google's measurements, not independent device tests [7].
The context budget sets the size of every indexing job. The model card gives one 8,192-token window shared across modalities, good for about 5.5 minutes of audio, 29 images or 58 video frames when only one is used [3]. Mixed inputs draw on the same budget, and video is sampled at one frame per second by default [3]. At that rate a single window covers 58 seconds of video [16]. One frame per second is fine for finding the birthday cake and less fine for finding the moment it hit the floor. An hour of audio needs at least 11 windows [17].
Matryoshka truncation lets developers cut vectors from 768 dimensions to 512, 256 or 128 [10]. Google's "up to 6x" storage saving is 768 divided by 128 [10][24]. The model card shows the quality cost. At 128 dimensions the multimodal benchmark score falls from 59.01 to 45.65, and the code score falls from 78.68 to 71.41 [18]. The loss is 13.36 points, about 23%, on mixed media against 7.27 points, about 9%, on code [26][27]. Google recommends 128 dimensions mainly for text-only workloads [18].
Code retrieval is where the release moves most against the first EmbeddingGemma. Multilingual MTEB went from 61.15 to 61.36, a 0.21-point change against a 9.92-point gain on MTEB Code [12][25]. For text-only retrieval the new model is close to a like-for-like swap. For the code gain to carry over, a team's queries and repository would need to resemble the benchmark's. These are Google's figures on Google's evaluation setup, and RuntimeWire notes they do not establish how the model performs on any private collection of photos, recordings, documents or code [19].
For on-device RAG, EmbeddingGemma 2 shares Gemma 4's text tokenizer and audio encoder, and Google says the two run together with a lower combined memory footprint [20]. Google did not publish the combined figure. The serving list is broad, covering llama.cpp, Ollama, vLLM, MLX and sentence-transformers among others [13]. The first EmbeddingGemma passed 20 million downloads [21]. On figures Google measured itself, the release supports the claim that one Apache-licensed model can do cross-modal search on a phone [2][6][7].
What to watch
- Independent RAM and latency measurements on non-Pixel phones, which would test whether Google's 191MB and 567MB figures hold on other hardware.
- Multimodal retrieval scores at 512 and 256 dimensions, the settings a mixed-media index would realistically use.
- Whether the planned ML Kit integration lets apps load the vision and audio encoders separately from the 270M text core.