On October 6, Google released EmbeddingGemma 2, an open model that lets photos and audio be searched with the same mechanism as text. On top of the text and code support of the original EmbeddingGemma, it adds images, video and audio, converting different types of data into a shared vector representation. The full model has 740 million parameters, and for text-only use you can load just the 270-million-parameter text portion. It is licensed under Apache 2.0.
Code retrieval scores improved substantially over the first generation, while the overall multilingual text score was nearly flat. Compared with competitors such as Qwen and Jina, the reasons to choose one model over another depend on what kind of data you want to search, how long the inputs are, and how much of the load you want to keep off the device.
740 million parameters, load only the capabilities you need
EmbeddingGemma 2 combines a 270-million-parameter text model with a 170-million-parameter image encoder and a 300-million-parameter audio encoder. Text plus images comes to about 440 million parameters, and text plus audio to about 570 million, so you can avoid loading the parts for modalities you don't need into memory. Video is processed as frames through the image encoder.
An embedding model does not generate answer text. It produces a sequence of numbers called a "vector" that represents the content of text, images, audio and so on. The content you want to search for and the stored data are each converted to vectors, and the closest matches are retrieved.
This means you can search by semantic similarity even when words don't match exactly. For example, you can build a system where you enter "shoes that don't slip on rainy mountain trails" and find product photos that fit.
Previously, it was already possible to generate captions from photos or transcribe audio and then pass the results to a text embedding model. With EmbeddingGemma 2, the image and audio encoders align what they extract into the same 768-dimensional vector space as text. This reduces the steps needed to generate intermediate captions or transcripts, and it makes it possible to search not only spoken content but also the images themselves and ambient sounds. Inputs that combine multiple formats, such as text and an image, can also be converted into a single vector.
However, finding relevant material and reading that material to write an answer are separate tasks. In retrieval-augmented generation (RAG), EmbeddingGemma 2 finds relevant data, and a generative model such as Gemma 4 produces the answer based on it. Google's published code search demo likewise pairs EmbeddingGemma 2 for retrieval with Gemma 4 on the agent side.
As an example of running a quantized model on the Pixel 11 Pro, Google says it runs with about 191MB of active RAM for text alone, and about 567MB with all modalities loaded. These figures apply to a specific device and quantization method, and they do not represent the app's total memory use, which would include the search index, the app itself and temporary working space during processing.
The published benchmarks are also for the full-precision model. It cannot be said that a small quantized on-device configuration would produce exactly the same scores.
Code search improves sharply; multilingual text is nearly flat
The clearest difference from the original EmbeddingGemma is in code retrieval. In the model card Google published, the MTEB code score rose from 68.76 to 78.68. The multilingual text score, by contrast, increased only slightly, from 61.15 to 61.36.
| Google-reported evaluation | EmbeddingGemma (original) | EmbeddingGemma 2 | Difference |
|---|---|---|---|
| MTEB Multilingual v2, task average | 61.15 | 61.36 | +0.21 |
| MTEB Code v1, NDCG@10 task average | 68.76 | 78.68 | +9.92 |
The source is the EmbeddingGemma 2 model card. The second-generation values are for full precision at 768 dimensions.
NDCG@10 measures how well a system ranks highly relevant candidates within its top 10 search results. It does not measure the ability to generate code or fix bugs.
Dividing the 9.92-point gain by the original 68.76 gives a relative improvement of about 14.4%. Google's "roughly 14% improvement" refers to this ratio of benchmark scores. It does not mean search is 14% faster, nor that the model can fix 14% more code defects. The improvement is meaningful first of all for uses such as describing a process in natural language and finding the related code.
In the code search comparison chart Google published, EmbeddingGemma 2 outperforms Qwen3-Embedding-0.6B and pplx-embed-v1-4b but falls short of Qwen3-Embedding-8B. It is notable that a model with a text portion of just 270 million parameters can compete with larger ones, but it has not become the top performer when models up to the 8B class are included.
This ranking also reflects benchmark results published by Google; it is not a comparison of on-device power consumption or search time under the same conditions.
Registration of the results with third-party benchmark infrastructure is still in progress. As of October 7, both the PR to register the model in MTEB and the PR submitting the evaluation results had not been merged.
A Google developer has said that the exact composition of the instructions used in evaluation cannot be disclosed, but that the instructions themselves were built from the prompts already provided, and that the results can be reproduced using MTEB and SentenceTransformers. On the MTEB side, however, there have been calls to make the instruction templates and implementation method clearer in order to ensure reproducibility.
The current figures are therefore not scores that have been found to be wrong, but neither are they values whose registration and verification MTEB has formally completed. It is reasonable to treat them as evaluation values published by Google and to check the registration results as they come in.
Qwen3-Embedding-0.6B is a model of about 600 million parameters aimed at text and code, supporting inputs of up to 32K tokens and output of up to 1024 dimensions. It can process longer inputs at once than EmbeddingGemma 2's 8K, and the same series also includes larger 4B and 8B models.
For splitting code into small units and indexing it on the device, EmbeddingGemma 2's small size can be put to good use. On the other hand, if you want to process long documents or large code files in one pass, Qwen's 32K input length becomes important.
In multilingual text, the difference from the original is only 0.21 points, so it is hard to say that retrieval accuracy itself improved significantly. The original was a text-only model of about 300 million parameters with a maximum input of 2K and 768-dimensional output.
The second generation shrinks the text portion to 270 million parameters while extending the input length to 8K and adding support for images, video and audio. Its significance lies less in greatly raising the multilingual score than in advancing miniaturization and multimodality at the same time.
Support for more than 100 languages also does not guarantee the retrieval accuracy needed for Japanese business documents. When adopting it, accuracy should be checked using the Japanese data you will actually handle.
For images and audio, rankings against competing models flip
On the image benchmark MIEB Lite, EmbeddingGemma 2's task-type average score was 64.64. In Google's comparison chart, it outperforms the base and so400m versions of SigLIP, which handle images and text, as well as jina-embeddings-v5-omni-small and BidirLM-Omni-2.5B-Embedding. LCO-Embedding-Omni-3B sits higher still.
Here, the score in the image domain needs to be separated from the model's overall capability.
For example, jina-embeddings-v5-omni-small has about 1.7 billion parameters in total and handles not only text but also images, video, audio and PDFs. EmbeddingGemma 2's 740 million parameters is less than half of that, but the image retrieval score alone does not determine which model is better across all uses.
On the audio benchmark MAEB, EmbeddingGemma 2 recorded 49.39. In Google's comparison chart it outperforms larger_clap_general, MuQ-MuLan-large, Qwen2-Audio-7B, Qwen2.5-Omni-3B and others, but falls short of jina-embeddings-v5-omni-nano, BidirLM and LCO-Embedding-Omni-7B.
Even though it beats Jina's omni-small on images, Jina's omni-nano ranks higher on audio. The competing models and the rankings change depending on the type of data being searched.
MAEB is an evaluation that includes not only speech but also music, environmental sounds and retrieval across audio and text. The MAEB preprint, published in February 2026, covers 30 tasks spanning more than 100 languages and also reports the strengths and weaknesses of individual models, such as models that are strong on environmental sounds but weak on multilingual speech.
For that reason, the score of 49.39 cannot be read as "49.39% accuracy in audio retrieval."
Google also published results on a different benchmark, MMEB v2: 57.28 for images, 67.84 for visual documents and 50.67 for video. However, it uses Hit@1 for images and video and NDCG@5 for visual documents, so the evaluation method also differs from MIEB Lite.
A high score in photo retrieval does not mean the model is equally strong at document retrieval that includes charts and figures, or at finding a specific moment within a video.
When actually choosing a model, you first need to decide whether you want to find what appears in a photo, what is being said in audio, or what a chart in a document shows.
Qwen, Jina and Gemini are suited to different uses
Qwen3-VL-Embedding is a retrieval model that converts text, images and video into a common vector. The 2B version supports 32K input and output of up to 2048 dimensions, and the 8B version can use up to 4096 dimensions. It can process longer inputs than EmbeddingGemma 2, but the official specifications do not include audio input.
Lining up the specifications of the main models shows that even among "embedding models," the retrieval systems they are designed for differ.
| Model | Reported parameters | Main inputs | Max input window | Default / max output dimensions |
|---|---|---|---|---|
| EmbeddingGemma 2 | 740M total, 270M text only | Text, images, video, audio | 8,192 tokens shared | 768, can be shortened to 128 |
| Qwen3-Embedding-0.6B | 0.6B | Text, code | 32K | Up to 1024 |
| Qwen3-VL-Embedding-2B | 2B | Text, images, video | 32K | Up to 2048 |
| jina-embeddings-v5-omni-small | About 1.7B | Text, images, video, audio, PDF | 32,768 | 1024, can be shortened to 32 |
| Gemini Embedding 2 | Not disclosed | Text, images, video, audio, PDF | 8,192 | Up to 3072 |
Specifications are from each developer's published materials. Because the way tokens are counted and the way images and video are processed differ from model to model, 32K support does not mean a model can always read four times as much Japanese text as an 8K model. The number of output vector dimensions also does not, by itself, determine the ranking of retrieval accuracy.
Qwen3-VL also publishes high scores on its own MMEB v2 evaluation, but it explicitly states that it updated the data used to evaluate visual documents. It is not a comparison with Google's published evaluation that aligns everything down to how inputs are fed to the model and the evaluation conditions.
Rather than simply subtracting benchmark numbers presented by different developers, narrowing candidates by use case, such as wanting to process long material at once or wanting to put audio into the same index, is more useful in actual design.
Jina's omni-small has the advantage of making it easy to reuse existing text indexes. The company says it shares a text vector space with jina-embeddings-v5-text-small, so an index already built with the text model can also be searched from images and audio.
If you have already built a large index with Jina's text model, you may be able to avoid re-converting all existing data in order to go multimodal.
On the other hand, Jina's published models are under the CC BY-NC 4.0 license, and for commercial use the company directs users to inquire separately. EmbeddingGemma 2 and the Qwen models compared here are released under Apache 2.0, so the conditions for building them into your own products also differ.
Even if the ability to download the weights is the same, commercial-use terms are not necessarily the same.
Gemini Embedding 2 has a similar name but is delivered differently. It is a multimodal embedding model used through the Gemini API, with execution of the model itself left to Google's service. It can generate vectors of up to 3072 dimensions, which can also be changed to 768 or 1536 dimensions, among others.
EmbeddingGemma 2, which is meant for searching photos and recordings on a device without sending them over the network, and Gemini Embedding 2, which uses a cloud API to generate large-scale indexes, assume fundamentally different modes of operation.
The amount of input they accept is also not the same. Gemini Embedding 2 handles up to 180 seconds of audio and up to 120 seconds of video, sampling video at up to 32 frames.
In EmbeddingGemma 2, text, images, audio and other inputs share a common 8,192-token window. Audio is counted at 25 tokens per second, and an image at the default setting is counted as 280 tokens.
For a standalone input with no accompanying text, simple arithmetic gives about 327 seconds of audio or the equivalent of about 29 images. For video, the token count alone would correspond to about 58 frames, but in the distributed input processing settings, the default is one frame per second with a maximum of 32 frames, and anything beyond that is sampled evenly.
When making long recordings or videos searchable, design decisions matter as well as the model's maximum input, such as what units to split them into and which scenes to keep in the index.
Also, even models that output vectors of the same 768 dimensions cannot necessarily have their vectors compared directly. Different models mean the space the numbers represent is itself different.
Google likewise explains that migrating from the first-generation gemini-embedding-001 to Gemini Embedding 2 requires re-converting existing data into embeddings with the new model. Cases like Jina that explicitly state compatibility with an existing model are specific designs, and there is no general rule that an index can be reused as is if the output dimensions match.
Shortening vectors changes both storage size and search accuracy
EmbeddingGemma 2 can shorten its standard 768-dimensional vectors to 512, 256 or 128 dimensions. This is possible because it uses Matryoshka Representation Learning (MRL), which trains the model so that the leading dimensions carry the most important information.
Shorter vectors reduce storage size and the computation needed at search time, but the effect on search accuracy varies greatly by use.
| Output dimensions | Vector component size ratio | Multilingual MTEB v2 | Code MTEB v1 | MMEB v2 overall | Audio MSEB retrieval |
|---|---|---|---|---|---|
| 768 | 1 | 61.36 | 78.68 | 59.01 | 69.54 |
| 512 | 2/3 | 61.17 | 77.24 | 58.38 | 69.18 |
| 256 | 1/3 | 60.41 | 76.18 | 56.24 | 66.76 |
| 128 | 1/6 | 57.89 | 71.41 | 45.65 | 56.71 |
The source is the EmbeddingGemma 2 evaluation by dimension, published on October 6. These are results from changing only the number of output dimensions of vectors from the same full-precision model. Because the evaluation metrics differ by column, what should be compared is the change within each column as dimensions are shortened.
In Google's published figures, shortening from 768 to 128 dimensions drops the overall MMEB v2 score from 59.01 to 45.65. The relative decline is about 22.6%, calculated as (59.01 − 45.65) ÷ 59.01 × 100.
This does not mean accuracy falls by 22.6 points. It is a relative change in the benchmark score.
In code retrieval, by contrast, the score goes from 78.68 to 71.41, and the same calculation gives a decline of only about 9.2%. This shows that the impact of shortening vectors differs greatly depending on what is being searched.
At 256 dimensions, the overall MMEB v2 score stays at 56.24 and code at 76.18. It cuts the number of stored vector elements to one third while retaining more retrieval performance than at 128 dimensions.
It is better to consider how far to cut dimensions separately for building a large text or code index to roughly narrow down candidates first, and for precisely finding target data in photos or video.
The effect on storage can be calculated simply. If you store 1 million vectors in float32 at 4 bytes per element, the vector components alone come to 3.072GB at 768 dimensions, 1.024GB at 256 dimensions and 0.512GB at 128 dimensions.
The calculation is 1 million vectors × number of dimensions × 4 bytes, with GB in decimal notation.
However, this is only the size of the vectors themselves. It does not include the index structures used for search, metadata, or the original photos, audio and documents. Even shrinking from 768 to 128 dimensions therefore does not cut the storage of the whole database or app to one sixth.
Quantization, which makes the model's weights themselves smaller, is also a separate process from reducing the number of dimensions of the generated vectors.
In addition, when vectors are shortened, L2 normalization must be redone, and the number of dimensions on the search query side and the stored data side must match. The model card advises using bfloat16 or float32 for normal inference and asks users to avoid float16, which can produce NaN values or degrade retrieval quality.
If you are building an app that searches photos and recordings together on a device, a realistic approach is to first check the retrieval accuracy you need at 768 or 256 dimensions, then decide the number of dimensions by balancing against index size.
On the other hand, if you have requirements such as processing long inputs at once or reusing an index you have already built, the designs of Qwen and Jina are also candidates.
The distinguishing feature of EmbeddingGemma 2 is not simply that it posts the highest benchmark scores. It is that it brings text, images, video and audio into a single search foundation, loads only the capabilities you need, and lets you adjust the vector size to match the device's storage and the accuracy you require.
