Google DeepMind
EmbeddingGemma 2
One compact model puts different media into a shared search space

Key facts
- 740M parameters
- Full model
- 270M parameters
- Text and code
- 768 values
- Vector size
- Apache 2.0
- Licence
EmbeddingGemma 2 turns words, pictures, video and sound into lists of numbers that software can compare. That lets one search find related content across different media, including on a phone or laptop.
Listen to the model guide
Podcast: Searching across different media
EmbeddingGemma 2 turns text, code, pictures, video and sound into lists of numbers that software can compare. Google released the open model on 6 October 2026, with downloadable weights under Apache 2.0.
Those numerical lists are called embeddings. Content with a similar meaning can sit close together in the model’s shared search space, so a written question can help find a relevant photograph or recording.
How does an embedding help you search?
EmbeddingGemma 2 represents each item with 768 numerical values by default. A search application embeds the query, compares it with stored items and retrieves the closest matches. Sentence Transformers’ semantic-search guide explains the query-and-document comparison used in this kind of retrieval.
Google’s model places text, code, images, video and audio in a shared space. A query about waves breaking on a beach can therefore be compared with a beach photograph, a clip or an audio recording.
The retrieved items can support a search interface or supply evidence to a separate answer-writing model. The application controls which files are indexed and which results a user can access. That access control stays important when a collection contains private documents.
How large is the model you need to load?
Google’s developer guide describes a modular network: a 270-million-parameter text and code model, with optional vision and audio encoders.
| Loaded inputs | Parameters |
|---|---|
| Text and code | 270M |
| Text, images and video | 440M |
| Text and audio | 570M |
| All supported media | 740M |
The vision encoder adds 170 million parameters; the audio encoder adds 300 million. Each configuration projects into the same compatible embedding space. A text-only query can therefore be compared with an item indexed using the full multimodal model.
Google designed EmbeddingGemma 2 for consumer hardware, including phones and laptops. Actual memory use depends on the loaded encoders, numerical precision, input length and software runtime. Selecting the components the application uses keeps deployment requirements tied to its task.

How much space do the vectors take?
EmbeddingGemma 2 supports 768, 512, 256 and 128 dimensions. Google’s model card explains how developers can shorten the vector and normalise it again for similarity comparisons.
At the same numerical precision, 256 values use one-third of the storage occupied by 768 values. The following arithmetic describes the vector values themselves; a database also stores its index and metadata.
| Dimensions | Bytes at 32 bits per value | Relative vector storage |
|---|---|---|
| 768 | 3,072 | 100% |
| 256 | 1,024 | 33.3% |
| 128 | 512 | 16.7% |
Shorter vectors trade storage and search work against retrieval quality. Google’s tests show a larger multimodal quality drop at 128 dimensions, so a team should compare results on its own collection before choosing that setting. Queries and stored items need the same vector length.

How do you use it with your own files?
Google publishes the weights as google/embeddinggemma-2 on Hugging Face. Its developer guide supports Sentence Transformers 6.1.0 or later, with separate query and document instructions for retrieval.
All media share an 8,192-token input budget. A mixed input spends that budget across its text, images, video frames and audio. Google’s model card recommends matching the task prefixes, normalising shortened vectors and selecting a supported numerical precision.
In Google’s reported code-retrieval evaluation, EmbeddingGemma 2 scored 78.68, compared with 68.76 for its predecessor, on MTEB Code v1’s NDCG@10 metric. That metric measures the ranking of relevant results. These are the publisher’s reported results, rather than a guarantee for every codebase.
The Apache 2.0 licence permits use, modification and redistribution, including commercial use, subject to its terms. Local deployment gives an application a way to process its own collection on the chosen device; testing relevant queries establishes whether its retrieval is useful.
More in Large Language Models
All LLMs →- GoogleGemma 4the laptop model
- Mistral AIMistral Large 4A trillion-parameter model opens in public preview
- Reflection AIBeamA sparse model with an October open-weight release planned
- Google DeepMindGemini 4 Argonlonger reasoning, coding and cyber defence
- OpenAIGPT-5.6the flagship from 9 July to 3 September 2026
- OpenAIGPT-6 Astrathe new generation, rolling out from 3 September 2026