Gemini Embedding

Gemini Embedding 2

by Google DeepMind · Generally available

Gemini Embedding 2 is Google DeepMind’s multimodal embedding model for converting text, images, video, audio, and PDFs into compatible vectors. It supports an 8,192-token input limit, configurable dimensions from 128 to 3,072, Gemini API and Vertex AI access, modality-based pricing, and Batch processing. It is intended for retrieval and organization rather than conversation or content generation.

Embeddings Reasoning Coding
Gemini Embedding 2 is designed for applications that need to compare meaning across text and media rather than generate conversational answers. A text description can be matched against an image, a video segment, an audio recording, or a PDF because the model places supported inputs in a shared embedding space. That makes it particularly relevant to cross-modal search, multimodal retrieval-augmented generation, recommendation systems, classification, clustering, and content indexing.
Outputs

What Gemini Embedding 2 can produce

Embeddings
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

2/10 Reasoning
2/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Gemini Embedding
Model type Embedding
Context window 8K tokens
Release date 2026-03-10
Status Generally available
Model notes

Canonical Gemini API model ID is gemini-embedding-2. It maps text, images, video, audio, and PDFs into a unified embedding space and returns numerical embeddings rather than generated content. Input limit is 8,192 tokens overall. Output dimensions are configurable from 128 to 3,072, with 768, 1,536, and 3,072 recommended. Images support up to six PNG or JPEG files per request; audio supports MP3 and WAV up to 180 seconds; video supports MP4 and MOV up to 120 seconds; PDFs support one file up to six pages. Gemini Embedding 2 vectors are incompatible with gemini-embedding-001 vectors, so migration requires re-embedding existing data. General availability was announced on April 22, 2026 after the initial public preview release.

Cost

Model pricing

Input Standard paid tier: text $0.20 per 1M tokens; images $0.45 per 1M tokens or $0.00012 per image; audio $6.50 per 1M tokens or $0.00016 per second; video $12.00 per 1M tokens or $0.00079 per frame. Batch pricing is 50% lower.
Output No separate output-token price; the model returns embeddings. Output is included in the input-modality pricing structure.
Model guide

Gemini Embedding 2: Google’s Multimodal Vector Model for Cross-Modal Search

Gemini Embedding 2 is Google DeepMind’s natively multimodal embedding model. It converts text, images, video, audio, and PDF documents into compatible numerical vectors, allowing applications to search, classify, cluster, recommend, and retrieve information across different media types. It supports an 8,192-token input limit, configurable output dimensions from 128 to 3,072, Gemini API and Vertex AI access, and modality-based pricing rather than generated-text pricing.

What is Gemini Embedding 2?

Gemini Embedding 2 is Google DeepMind’s first natively multimodal embedding model. Instead of returning prose, an image, or another generated media file, it converts supported content into a numerical vector: a list of values that represents the content’s meaning in a form that software can compare.

The important distinction is that Gemini Embedding 2 is a retrieval and representation model, not a conversational model. For example, it can help an application find images related to a text query, identify video content matching a description, or retrieve a PDF that is semantically similar to an audio recording. A separate generative model is still needed to write an answer, summarize retrieved material, or create new media.

Google provides the model through the Gemini API and Vertex AI. Its stable Gemini API model identifier is gemini-embedding-2. Within Google’s model lineup, it serves a specialized role alongside generative Gemini models: it provides the vector representations used to organize and retrieve information, including information that spans several media types.

How its shared embedding space works

Traditional search systems often require separate pipelines for text, images, audio, and video. Gemini Embedding 2 is designed to reduce that separation. Content from different modalities is mapped into a unified vector space, where related meanings should be located near one another according to a similarity calculation.

In practical terms, a search application could embed a query such as “a red car driving through a snowy mountain road” and compare it with vectors created from photographs or video clips. The same general approach can be used to search a document collection using multimedia evidence, find audio recordings related to written descriptions, or connect a product image with relevant text documents.

The model also supports interleaved multimodal content. Text can be submitted together with an image or another supported input, and the model can produce an aggregated embedding for the combined content. This is useful when the meaning depends on both a caption and the media it describes.

Supported inputs and processing limits

Gemini Embedding 2 accepts text, images, audio, video, and PDF documents. The overall input limit is 8,192 tokens, shared across the content included in a request. The token limit is therefore not simply a promise that every supported file can be submitted at its maximum duration or size at the same time.

Input typeSupported detailsResearch-backed limit
TextText content for semantic representationUp to 8,192 tokens
ImagesPNG and JPEGUp to six images per request
AudioMP3 and WAVUp to 180 seconds
VideoMP4 and MOV; H.264, H.265, AV1, and VP9 codecsUp to 120 seconds
PDFPDF document inputOne PDF of up to six pages per request

These limits influence system design. Long documents may need to be divided into chunks before indexing. Longer videos may need to be processed in segments, and a content library may require multiple embedding requests for a single source item. Chunking and segmentation also make it possible to retrieve a more precise passage or media interval instead of returning an entire large file.

Embedding dimensions and storage trade-offs

The model supports configurable embedding dimensions from 128 to 3,072. Google recommends 768, 1,536, or 3,072 dimensions for common quality and storage trade-offs. Higher-dimensional vectors generally require more storage and more computational work during similarity searches, while smaller vectors can make large indexes cheaper and faster to operate.

Gemini Embedding 2 uses Matryoshka Representation Learning, which allows a larger representation to be truncated for more efficient storage and comparison. The model automatically normalizes embeddings generated at reduced dimensions. That simplifies cosine-similarity workflows because an application does not need to perform a separate normalization step when requesting dimensions below the default 3,072.

The appropriate dimension depends on the application’s accuracy, latency, and infrastructure requirements. A large archive with strict storage constraints may prefer a smaller representation, while a high-value retrieval system may choose a larger one. The supplied research does not provide benchmark results showing how quality changes at each dimension, so the dimension choice should be validated against the application’s own evaluation data.

What Gemini Embedding 2 is best used for

Cross-modal search is the model’s clearest differentiator. A text query can be compared with vectors for images, videos, audio, or PDFs in the same vector space. This can support a video library searched with natural-language descriptions, an image archive searched with written prompts, or a media repository where text and non-text evidence are indexed together.

Multimodal retrieval-augmented generation

Gemini Embedding 2 can provide the retrieval layer for a multimodal retrieval-augmented generation, or RAG, system. The application first embeds its source material, stores the vectors in a search index, and retrieves records that resemble a user’s query. A separate generative model can then receive the selected text, images, audio, video, or document context and produce an answer or summary.

This separation is useful because the embedding model focuses on finding relevant material, while the generative model focuses on explaining it. It also means that Gemini Embedding 2 should not be evaluated as though it were a chatbot: it does not directly answer questions or provide natural-language explanations.

Classification, clustering, and recommendations

Vectors can be used as features for similarity-based classification, duplicate detection, content clustering, and recommendation systems. For example, a service could group mixed media by topic, find visually or semantically similar items, identify potentially duplicate content, or recommend documents and media related to a user’s current item.

These uses do not require the model to generate output in the usual sense. The returned embeddings are numerical data that an application stores and processes with a vector database or another similarity-search system.

Pricing and API availability

Gemini Embedding 2 is available through the Gemini API and Vertex AI. Its pricing is based on the input modality because the model returns embeddings rather than generated text. On the standard paid tier, the supplied pricing is:

  • Text: $0.20 per 1 million tokens.
  • Images: $0.45 per 1 million tokens, with an equivalent listed rate of $0.00012 per image.
  • Audio: $6.50 per 1 million tokens, with an equivalent listed rate of $0.00016 per second.
  • Video: $12.00 per 1 million tokens, with an equivalent listed rate of $0.00079 per frame.

Batch processing is available at 50% of the standard input rates. Actual costs can vary with modality, tokenization, video frame processing, and whether the workload runs through the Gemini API or Vertex AI. The per-image, per-second, and per-frame equivalents should therefore be treated as practical listed rates rather than a guarantee that every request will be billed solely by that unit.

Capabilities and non-capabilities

Gemini Embedding 2 supports multimodal input and embedding output. It does not provide direct text, image, video, audio, music, or action output. It also is not described as a reasoning or coding model, and the supplied model data lists tool use, streaming, fine-tuning, caching, and JSON mode as unsupported or unavailable for this model.

That does not prevent it from being used inside a sophisticated application. A developer can combine it with a vector database, a generative model, metadata filters, and application logic. The distinction is that those capabilities belong to the surrounding system, not to Gemini Embedding 2 itself.

Its main speed and cost advantage comes from specialization and configurable vector size rather than from generating long responses. Text embeddings are substantially less expensive than the listed audio and video rates, while media inputs may require more processing. Batch pricing can reduce the cost of large offline indexing jobs, but interactive media search may need to balance retrieval quality, latency, and per-request expense.

Migration from Gemini Embedding 001

Gemini Embedding 2 uses an embedding space that is incompatible with gemini-embedding-001. Existing vectors from the older model should not be directly compared with vectors produced by Gemini Embedding 2. A migration therefore requires re-embedding the indexed content and replacing the stored vectors before switching similarity searches to the new model.

The request design also differs. Unlike gemini-embedding-001, Gemini Embedding 2 does not use the older task_type parameter. For text-only retrieval and similarity tasks, Google recommends expressing the task instruction in the prompt. Applications should review their request construction and evaluation process rather than treating the migration as a simple model-name change.

Limitations to plan for

The model’s most important limitation is its purpose: it produces representations, not explanations. It cannot independently summarize a retrieved PDF, answer a user’s question, write code, or generate a visual result. Those tasks require a separate model after retrieval or analysis.

Input limits can also affect recall and indexing strategy. A PDF is limited to one file of up to six pages per request, audio is limited to 180 seconds, and video is limited to 120 seconds. Large or long sources must be divided into smaller units, which introduces an application-level choice about chunk boundaries, overlap, timestamps, and metadata.

There is no supplied benchmark evidence for exact retrieval quality, latency under different workloads, or the best embedding dimension for every domain. Those factors should be measured using representative queries and documents. In particular, a system that combines several modalities should test whether its chosen aggregation method retrieves the right media, rather than assuming that a shared vector space removes all domain-specific tuning.

When to choose Gemini Embedding 2

Choose Gemini Embedding 2 when the application needs one embedding workflow for multiple media types, especially when a query in one modality must retrieve content in another. It is a strong fit for cross-modal search, mixed-media RAG, recommendations across text and media, and large collections that need semantic organization.

It is also a reasonable choice when configurable dimensions matter. A team can select a smaller vector representation to reduce storage and search costs or use a larger representation when its evaluation justifies the additional expense. Batch processing is useful for indexing existing collections rather than embedding every item interactively.

A text-only embedding option may be more appropriate when the data is exclusively textual and the additional multimodal capability provides no practical value. A generative Gemini model is more appropriate when the required result is an answer, summary, explanation, generated code, or new media. For applications already using gemini-embedding-001, Gemini Embedding 2 may provide a broader modality range, but the incompatibility of the vector spaces means that migration includes a full re-embedding step.

Overall, Gemini Embedding 2 is best understood as infrastructure for finding and organizing information. Its value comes from connecting different kinds of content in a common semantic index, while the surrounding application and other models handle storage, retrieval policy, reasoning, and user-facing generation.


Answers to Frequently Asked Questions

Does Gemini Embedding 2 generate answers or summaries?
No. Gemini Embedding 2 produces numerical embeddings and does not directly generate text, summaries, code, images, audio, video, or other media. A separate generative model is required to answer questions or summarize the content retrieved with its embeddings.
How does Gemini Embedding 2 differ from Gemini Embedding 001?
Gemini Embedding 2 uses an embedding space that is incompatible with gemini-embedding-001, so existing vectors cannot be directly compared with new vectors. Migration requires re-embedding indexed content, replacing stored vectors, and updating request construction because Gemini Embedding 2 does not use the older task_type parameter.
Can Gemini Embedding 2 search across different media types?
Yes. Gemini Embedding 2 maps different modalities into a shared vector space, allowing a text query to retrieve related images, videos, audio recordings, or PDFs. It also supports interleaved multimodal content, such as text combined with an image, to create an aggregated embedding.
What input types and limits does Gemini Embedding 2 support?
Gemini Embedding 2 supports text, PNG and JPEG images, MP3 and WAV audio, MP4 and MOV video, and PDF documents. Requests support up to 8,192 tokens overall, six images, 180 seconds of audio, 120 seconds of video, or one PDF of up to six pages, depending on the content submitted.
What is Gemini Embedding 2 used for?
Gemini Embedding 2 converts text, images, audio, video, and PDFs into numerical vectors for semantic retrieval and similarity comparison. It is designed for cross-modal search, multimodal RAG, classification, clustering, duplicate detection, and recommendation systems, rather than generating answers or media.


Sources 7
Provider

About Google DeepMind