BuildNot yet confirmed elsewhere1 publisher3 min readPublished Updated
EmbeddingGemma 2 maps five modalities into one 768-dimension vector space
Google DeepMind released EmbeddingGemma 2, a 740M-parameter open model that runs on a phone and embeds text, code, images, video and audio in one space. Teams running a separate embedder per modality can consolidate on it if Google's reported benchmark numbers hold on their own data.
The Engineer · Build desk
What happened
- One checkpoint holds every encoder, and a pipeline loads only what it uses: 270M parameters for text and code, 440M with vision added, 570M with audio.
- The model shares its text tokenizer and audio encoder architecture with Gemma 4, so pairing the two needs less memory than two independent models.
- Weights are published on Hugging Face and Kaggle under the Apache 2.0 license and are cleared for commercial use.
- The first EmbeddingGemma, released in 2025, handled text only and passed 20 million downloads.
Why it matters
- decision A retrieval test before retiring a specialist encoder has to use the quantized build that will ship on the device, not the full-precision weights Google scored.
- contradiction Vector storage sized from the writeup's 6x figure at 256 dimensions will come out at twice the planned footprint in bfloat16.
- constraint A single index that mixes media has to stay wider than a text-only index, because multimodal recall degrades faster under truncation.
- exposure With no post-training alignment or output moderation in the model, filtering what retrieval surfaces falls to the team deploying it.
EmbeddingGemma 2 is built on Gemma 4's transformer backbone, with 24 layers, both grouped-query and multi-query attention, a vocabulary of 262,144 tokens, and outputs mean-pooled into a projection layer that takes them from 512 dimensions to 768 [13]. Every modality projects into those same 768 dimensions. A text query and an audio clip are then compared with plain cosine similarity [4]. Inputs can be interleaved, text and images in one pass, with placeholder tokens marking each media item's position [7]. All five modalities draw on one 8,192-token context window, and each spends it at a fixed rate [6]. The first model's window was 2,048 tokens [21].
The packaging deserves credit. The module sizes add up exactly: a 130M backbone and a 140M embedder for text, a 170M vision encoder, a 300M audio encoder, 740M in total [17]. Audio is the largest single piece, about 41% of the full model [18]. An app that only does text-to-image retrieval never loads it [5]. On the Pixel 11 Pro figures, the vision and audio encoders together add about 376MB of active RAM over the text-only build [19].
Matryoshka Representation Learning is what makes the vectors cuttable. The first N dimensions of any output are a valid embedding on their own, so 768 dimensions can be truncated to 512, 256 or 128 without retraining [15]. The dev.to writeup's storage figures hold up: one million 768-dimension vectors in bfloat16 come to roughly 1.5GB, and about 250MB at 128 dimensions [23]. That is 1,536 bytes a vector at full width [24]. One line in its quality table conflicts with them. The table, attributed to the model card, puts 256 dimensions at about 95% of full multimodal quality with a ~6x storage cut [22]. At two bytes a dimension, 256 dimensions is 512 bytes a vector, or 512MB per million, a 3x cut [24]. The 6x figure fits the 128-dimension row, or a float32 baseline [25][26]. At 128 dimensions, quality holds near 90% on text and code and falls to about 75% on multimodal retrieval [22].
Google reports its scores on the full-precision model [8]. The Pixel footprints are for a quantized build [10], and the writeup does not include retrieval scores for that build. Code shows the clearest gain: MTEB Code rose 9.92 points over EmbeddingGemma 1, to 78.68 [20][8]. The media scores sit lower, 57.28 on MMEB v2 image retrieval, 50.67 on video retrieval and 49.39 on MAEB audio [9], but those are separate benchmarks on separate scales. According to Google's reported numbers, the model beats some specialist models more than twice its size on audio and visual benchmarks [27]. For any of this to transfer, your queries have to resemble the benchmark queries, and quantization has to cost little on your hardware.
Integration cost is low. The model runs in transformers, sentence-transformers 6.1.0 and later, vLLM, llama.cpp, SGLang and Ollama [11]. Browsers get it through transformers.js and WebGPU, phones through MediaPipe and LiteRT [11]. The case for collapsing per-modality models into this one checkpoint is sound on architecture and unproven on quality outside Google's own tables. In my context, a team running separate text and image encoders, I'd load the text-and-vision configuration, embed a sample of the corpus at 256 dimensions, and compare recall against the current pipeline before retiring anything.
What to watch
- Independent evaluations of the quantized on-device build against Google's full-precision MMEB, MAEB and video retrieval scores.
- A corrected model card table, or a statement of whether the ~6x figure at 256 dimensions is measured against a float32 baseline.
- Published per-modality token rates for the 8,192-token window, since they set how much video or audio fits in one embedding call.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption20
- Hype gap+25
- Incentives60
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Google DeepMind released EmbeddingGemma 2 on October 6, 2026.
- [2]
EmbeddingGemma 2 is a single 740M-parameter open model that maps text, code, images, video and audio into one shared vector space, designed to run entirely on-device.
- [3]
The original EmbeddingGemma (2025) was text-only and accumulated over 20 million downloads.
- [4]
All modalities project into the same 768-dimensional embedding space, so a text query and an audio clip can be compared directly by cosine similarity.
- [5]
From a single checkpoint you load only the encoders you need: text and code 270M parameters (130M transformer backbone + 140M embedder); text + vision 440M (adds 170M vision encoder); text + audio 570M (adds 300M audio encoder); full multimodal 740M. An app doing only text-to-image retrieval does not pay the memory cost of the audio encoder.
- [6]
All five modalities share an 8,192-token context window, four times larger than EmbeddingGemma 1, and each modality consumes that budget at fixed rates.
- [7]
Inputs can be interleaved, text and images mixed in one pass, with placeholder tokens marking each media item's position.
- [8]
Google reports, on the full-precision model, MTEB Code of 78.68, up from 68.76 in EmbeddingGemma 1.
- [9]
Other Google-reported scores: MTEB multilingual (v2) 61.36, MIEB lite 64.64, MMEB v2 image retrieval 57.28, visual-document retrieval 67.84, video retrieval 50.67, MSEB sound retrieval 69.54, MAEB audio 49.39.
- [10]
With quantization on a Pixel 11 Pro, text-only uses about 191MB of active RAM and full multimodal about 567MB.
- [11]
The model runs in the browser via transformers.js and WebGPU, on mobile and edge via MediaPipe and LiteRT, and server-side with transformers, sentence-transformers 6.1.0+, vLLM, llama.cpp, SGLang, Ollama and LMStudio; vector storage integrations include Qdrant.
- [12]
Weights are available on Hugging Face and Kaggle under the Apache 2.0 license and are commercially usable.
- [13]
EmbeddingGemma 2 is built on the Gemma 4 transformer backbone: 24 layers, grouped-query and multi-query attention, a 262,144-token vocabulary, and mean pooling with a projection layer from 512 to 768 dimensions.
- [14]
The model shares its text tokenizer and audio encoder architecture with Gemma 4; pairing EmbeddingGemma 2 with Gemma 4 gives a lower combined memory footprint than running two independent models.
- [15]
EmbeddingGemma 2 uses Matryoshka Representation Learning, making the first N dimensions of any output a valid embedding; embeddings can be truncated from 768 to 512, 256 or 128 dimensions without retraining.
- [16]
The model card states EmbeddingGemma 2 has no safety tuning: it is a pre-trained embedding model with no post-training alignment or output-level moderation, and safety mitigations are limited to pre-training data filtering such as CSAM filtering and automated PII removal.
- [17]
The component sizes sum to the full model: 130M + 140M + 170M + 300M = 740M parameters.
- [18]
The audio encoder is about 41% of the full 740M-parameter model.
- [19]
On the Pixel 11 Pro, the vision and audio encoders add about 376MB of active RAM over the text-only configuration.
- [20]
MTEB Code improved by 9.92 points from EmbeddingGemma 1 to EmbeddingGemma 2.
- [21]
EmbeddingGemma 1's context window was 2,048 tokens.
- [22]
Per the model card as reported by the writeup: 256 dimensions gives about 95% of full quality on multimodal retrieval with about a 6x storage reduction; 128 dimensions gives about 90% on text/code but drops to about 75% on multimodal retrieval.
- [23]
Storing one million 768-dimensional vectors in bfloat16 takes roughly 1.5GB; at 128 dimensions it drops to about 250MB.
- [24]
In bfloat16 (2 bytes per dimension) a 768-dimension vector is 1,536 bytes and a 256-dimension vector is 512 bytes, so one million 256-dimension vectors take about 512MB, a 3x reduction from 768 dimensions.
- [25]
A 6x storage reduction corresponds to truncating from 768 to 128 dimensions, matching the writeup's 1.5GB to 250MB example.
- [26]
A 6x reduction at 256 dimensions holds only against a float32 768-dimension baseline.
- [27]
According to Google's reported numbers, the model outperforms some specialist models more than twice its size on audio and visual benchmarks.
ReportedInsufficientSource: Google's reported numbers, as relayed by the dev.to writeup2 sources— create a free account to open themView cited source
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toEmbeddingGemma 2: How Google's Open Multimodal Embedding Model Works
2 articles · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.