What is Gemini Embedding 2?
Gemini Embedding 2 is Google DeepMind’s first natively multimodal embedding model. Instead of returning prose, an image, or another generated media file, it converts supported content into a numerical vector: a list of values that represents the content’s meaning in a form that software can compare.
The important distinction is that Gemini Embedding 2 is a retrieval and representation model, not a conversational model. For example, it can help an application find images related to a text query, identify video content matching a description, or retrieve a PDF that is semantically similar to an audio recording. A separate generative model is still needed to write an answer, summarize retrieved material, or create new media.
Google provides the model through the Gemini API and Vertex AI. Its stable Gemini API model identifier is gemini-embedding-2. Within Google’s model lineup, it serves a specialized role alongside generative Gemini models: it provides the vector representations used to organize and retrieve information, including information that spans several media types.
How its shared embedding space works
Traditional search systems often require separate pipelines for text, images, audio, and video. Gemini Embedding 2 is designed to reduce that separation. Content from different modalities is mapped into a unified vector space, where related meanings should be located near one another according to a similarity calculation.
In practical terms, a search application could embed a query such as “a red car driving through a snowy mountain road” and compare it with vectors created from photographs or video clips. The same general approach can be used to search a document collection using multimedia evidence, find audio recordings related to written descriptions, or connect a product image with relevant text documents.
The model also supports interleaved multimodal content. Text can be submitted together with an image or another supported input, and the model can produce an aggregated embedding for the combined content. This is useful when the meaning depends on both a caption and the media it describes.
Supported inputs and processing limits
Gemini Embedding 2 accepts text, images, audio, video, and PDF documents. The overall input limit is 8,192 tokens, shared across the content included in a request. The token limit is therefore not simply a promise that every supported file can be submitted at its maximum duration or size at the same time.
| Input type | Supported details | Research-backed limit |
|---|---|---|
| Text | Text content for semantic representation | Up to 8,192 tokens |
| Images | PNG and JPEG | Up to six images per request |
| Audio | MP3 and WAV | Up to 180 seconds |
| Video | MP4 and MOV; H.264, H.265, AV1, and VP9 codecs | Up to 120 seconds |
| PDF document input | One PDF of up to six pages per request |
These limits influence system design. Long documents may need to be divided into chunks before indexing. Longer videos may need to be processed in segments, and a content library may require multiple embedding requests for a single source item. Chunking and segmentation also make it possible to retrieve a more precise passage or media interval instead of returning an entire large file.
Embedding dimensions and storage trade-offs
The model supports configurable embedding dimensions from 128 to 3,072. Google recommends 768, 1,536, or 3,072 dimensions for common quality and storage trade-offs. Higher-dimensional vectors generally require more storage and more computational work during similarity searches, while smaller vectors can make large indexes cheaper and faster to operate.
Gemini Embedding 2 uses Matryoshka Representation Learning, which allows a larger representation to be truncated for more efficient storage and comparison. The model automatically normalizes embeddings generated at reduced dimensions. That simplifies cosine-similarity workflows because an application does not need to perform a separate normalization step when requesting dimensions below the default 3,072.
The appropriate dimension depends on the application’s accuracy, latency, and infrastructure requirements. A large archive with strict storage constraints may prefer a smaller representation, while a high-value retrieval system may choose a larger one. The supplied research does not provide benchmark results showing how quality changes at each dimension, so the dimension choice should be validated against the application’s own evaluation data.
What Gemini Embedding 2 is best used for
Cross-modal semantic search
Cross-modal search is the model’s clearest differentiator. A text query can be compared with vectors for images, videos, audio, or PDFs in the same vector space. This can support a video library searched with natural-language descriptions, an image archive searched with written prompts, or a media repository where text and non-text evidence are indexed together.
Multimodal retrieval-augmented generation
Gemini Embedding 2 can provide the retrieval layer for a multimodal retrieval-augmented generation, or RAG, system. The application first embeds its source material, stores the vectors in a search index, and retrieves records that resemble a user’s query. A separate generative model can then receive the selected text, images, audio, video, or document context and produce an answer or summary.
This separation is useful because the embedding model focuses on finding relevant material, while the generative model focuses on explaining it. It also means that Gemini Embedding 2 should not be evaluated as though it were a chatbot: it does not directly answer questions or provide natural-language explanations.
Classification, clustering, and recommendations
Vectors can be used as features for similarity-based classification, duplicate detection, content clustering, and recommendation systems. For example, a service could group mixed media by topic, find visually or semantically similar items, identify potentially duplicate content, or recommend documents and media related to a user’s current item.
These uses do not require the model to generate output in the usual sense. The returned embeddings are numerical data that an application stores and processes with a vector database or another similarity-search system.
Pricing and API availability
Gemini Embedding 2 is available through the Gemini API and Vertex AI. Its pricing is based on the input modality because the model returns embeddings rather than generated text. On the standard paid tier, the supplied pricing is:
- Text: $0.20 per 1 million tokens.
- Images: $0.45 per 1 million tokens, with an equivalent listed rate of $0.00012 per image.
- Audio: $6.50 per 1 million tokens, with an equivalent listed rate of $0.00016 per second.
- Video: $12.00 per 1 million tokens, with an equivalent listed rate of $0.00079 per frame.
Batch processing is available at 50% of the standard input rates. Actual costs can vary with modality, tokenization, video frame processing, and whether the workload runs through the Gemini API or Vertex AI. The per-image, per-second, and per-frame equivalents should therefore be treated as practical listed rates rather than a guarantee that every request will be billed solely by that unit.
Capabilities and non-capabilities
Gemini Embedding 2 supports multimodal input and embedding output. It does not provide direct text, image, video, audio, music, or action output. It also is not described as a reasoning or coding model, and the supplied model data lists tool use, streaming, fine-tuning, caching, and JSON mode as unsupported or unavailable for this model.
That does not prevent it from being used inside a sophisticated application. A developer can combine it with a vector database, a generative model, metadata filters, and application logic. The distinction is that those capabilities belong to the surrounding system, not to Gemini Embedding 2 itself.
Its main speed and cost advantage comes from specialization and configurable vector size rather than from generating long responses. Text embeddings are substantially less expensive than the listed audio and video rates, while media inputs may require more processing. Batch pricing can reduce the cost of large offline indexing jobs, but interactive media search may need to balance retrieval quality, latency, and per-request expense.
Migration from Gemini Embedding 001
Gemini Embedding 2 uses an embedding space that is incompatible with gemini-embedding-001. Existing vectors from the older model should not be directly compared with vectors produced by Gemini Embedding 2. A migration therefore requires re-embedding the indexed content and replacing the stored vectors before switching similarity searches to the new model.
The request design also differs. Unlike gemini-embedding-001, Gemini Embedding 2 does not use the older task_type parameter. For text-only retrieval and similarity tasks, Google recommends expressing the task instruction in the prompt. Applications should review their request construction and evaluation process rather than treating the migration as a simple model-name change.
Limitations to plan for
The model’s most important limitation is its purpose: it produces representations, not explanations. It cannot independently summarize a retrieved PDF, answer a user’s question, write code, or generate a visual result. Those tasks require a separate model after retrieval or analysis.
Input limits can also affect recall and indexing strategy. A PDF is limited to one file of up to six pages per request, audio is limited to 180 seconds, and video is limited to 120 seconds. Large or long sources must be divided into smaller units, which introduces an application-level choice about chunk boundaries, overlap, timestamps, and metadata.
There is no supplied benchmark evidence for exact retrieval quality, latency under different workloads, or the best embedding dimension for every domain. Those factors should be measured using representative queries and documents. In particular, a system that combines several modalities should test whether its chosen aggregation method retrieves the right media, rather than assuming that a shared vector space removes all domain-specific tuning.
When to choose Gemini Embedding 2
Choose Gemini Embedding 2 when the application needs one embedding workflow for multiple media types, especially when a query in one modality must retrieve content in another. It is a strong fit for cross-modal search, mixed-media RAG, recommendations across text and media, and large collections that need semantic organization.
It is also a reasonable choice when configurable dimensions matter. A team can select a smaller vector representation to reduce storage and search costs or use a larger representation when its evaluation justifies the additional expense. Batch processing is useful for indexing existing collections rather than embedding every item interactively.
A text-only embedding option may be more appropriate when the data is exclusively textual and the additional multimodal capability provides no practical value. A generative Gemini model is more appropriate when the required result is an answer, summary, explanation, generated code, or new media. For applications already using gemini-embedding-001, Gemini Embedding 2 may provide a broader modality range, but the incompatibility of the vector spaces means that migration includes a full re-embedding step.
Overall, Gemini Embedding 2 is best understood as infrastructure for finding and organizing information. Its value comes from connecting different kinds of content in a common semantic index, while the surrounding application and other models handle storage, retrieval policy, reasoning, and user-facing generation.

