YFarmX logoYFarmX

Google DeepMind

EmbeddingGemma 2

One compact model puts different media into a shared search space

Released 6 October 20264 min readLarge Language Models

Editorial illustration of a phone organising a photograph, a film strip, a sound recording and a text card, headed EmbeddingGemma 2.

Key facts

740M parameters
Full model
270M parameters
Text and code
768 values
Vector size
Apache 2.0
Licence

EmbeddingGemma 2 turns words, pictures, video and sound into lists of numbers that software can compare. That lets one search find related content across different media, including on a phone or laptop.

Listen to the model guide

Podcast: Searching across different media

EmbeddingGemma 2 turns text, code, pictures, video and sound into lists of numbers that software can compare. Google released the open model on 6 October 2026, with downloadable weights under Apache 2.0.

Those numerical lists are called embeddings. Content with a similar meaning can sit close together in the model’s shared search space, so a written question can help find a relevant photograph or recording.

EmbeddingGemma 2 represents each item with 768 numerical values by default. A search application embeds the query, compares it with stored items and retrieves the closest matches. Sentence Transformers’ semantic-search guide explains the query-and-document comparison used in this kind of retrieval.

Google’s model places text, code, images, video and audio in a shared space. A query about waves breaking on a beach can therefore be compared with a beach photograph, a clip or an audio recording.

The retrieved items can support a search interface or supply evidence to a separate answer-writing model. The application controls which files are indexed and which results a user can access. That access control stays important when a collection contains private documents.

How large is the model you need to load?

Google’s developer guide describes a modular network: a 270-million-parameter text and code model, with optional vision and audio encoders.

Loaded inputs Parameters
Text and code 270M
Text, images and video 440M
Text and audio 570M
All supported media 740M

The vision encoder adds 170 million parameters; the audio encoder adds 300 million. Each configuration projects into the same compatible embedding space. A text-only query can therefore be compared with an item indexed using the full multimodal model.

Google designed EmbeddingGemma 2 for consumer hardware, including phones and laptops. Actual memory use depends on the loaded encoders, numerical precision, input length and software runtime. Selecting the components the application uses keeps deployment requirements tied to its task.

Modular model diagram showing a 270M text and code backbone, a 170M vision addition and a 300M audio addition, making 740M parameters in the complete model.
The text backbone and optional vision and audio encoders total 740 million parameters. Source: Google's developer guide.

How much space do the vectors take?

EmbeddingGemma 2 supports 768, 512, 256 and 128 dimensions. Google’s model card explains how developers can shorten the vector and normalise it again for similarity comparisons.

At the same numerical precision, 256 values use one-third of the storage occupied by 768 values. The following arithmetic describes the vector values themselves; a database also stores its index and metadata.

Dimensions Bytes at 32 bits per value Relative vector storage
768 3,072 100%
256 1,024 33.3%
128 512 16.7%

Shorter vectors trade storage and search work against retrieval quality. Google’s tests show a larger multimodal quality drop at 128 dimensions, so a team should compare results on its own collection before choosing that setting. Queries and stored items need the same vector length.

Three illustrated vector-value cards showing 768 values using 3072 bytes, 256 using 1024 bytes and 128 using 512 bytes, at 32 bits per value.
Vector-value storage scales with length at the same precision. These figures use 32 bits per value; database indexes and metadata add space.

How do you use it with your own files?

Google publishes the weights as google/embeddinggemma-2 on Hugging Face. Its developer guide supports Sentence Transformers 6.1.0 or later, with separate query and document instructions for retrieval.

All media share an 8,192-token input budget. A mixed input spends that budget across its text, images, video frames and audio. Google’s model card recommends matching the task prefixes, normalising shortened vectors and selecting a supported numerical precision.

In Google’s reported code-retrieval evaluation, EmbeddingGemma 2 scored 78.68, compared with 68.76 for its predecessor, on MTEB Code v1’s NDCG@10 metric. That metric measures the ranking of relevant results. These are the publisher’s reported results, rather than a guarantee for every codebase.

The Apache 2.0 licence permits use, modification and redistribution, including commercial use, subject to its terms. Local deployment gives an application a way to process its own collection on the chosen device; testing relevant queries establishes whether its retrieval is useful.