Gemini Embedding

Gemini Embedding

by Google DeepMind · Current, scheduled for shutdown on 2028-05-14

Google DeepMind’s Gemini Embedding model converts text into configurable vector representations for semantic search, RAG, classification, clustering, recommendations, and similarity matching. It supports up to 3,072 dimensions, accepts up to 2,048 input tokens, is available through the Gemini API and Vertex AI, and is scheduled to shut down on May 14, 2028.

Embeddings Reasoning Coding
Gemini Embedding is a specialized representation model rather than a conversational Gemini model. It turns text into numerical vectors that applications can compare to find similar meaning, retrieve relevant passages, organize documents, or power recommendation systems. Its configurable vector size makes it possible to balance retrieval quality against storage and search cost, but its text-only design and announced 2028 shutdown are important considerations when selecting it for a new system.
Outputs

What Gemini Embedding can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Gemini Embedding
Model type Other
Context window 2K tokens
Release date 2025-07-14
Status Current, scheduled for shutdown on 2028-05-14
Shutdown date 2028-05-14
Knowledge cutoff notes

A specific training knowledge cutoff is not published in the reviewed official model documentation. The model should be treated as an embedding model rather than a knowledge-answering model.

Model notes

The canonical API model ID is gemini-embedding-001. It accepts text and returns text embedding vectors rather than natural-language output. Output dimensionality is configurable from 128 to 3072, with 768, 1536, and 3072 recommended. The model has a 2048-token input limit. Google lists Gemini Embedding 2 as the recommended replacement and states that gemini-embedding-001 will shut down on May 14, 2028. Pricing in this record uses Vertex AI rates; Gemini API availability and pricing may vary by billing surface.

Cost

Model pricing

Input $0.00015 per 1,000 input tokens for online Vertex AI requests; $0.00012 per 1,000 input tokens for batch requests
Output No charge for embedding output
Model guide

Gemini Embedding: Google’s Text-Only Model for Semantic Search and RAG

Gemini Embedding is Google DeepMind’s text-only embedding model, identified as gemini-embedding-001. It converts text into configurable-dimensional numerical vectors for semantic search, retrieval-augmented generation, classification, clustering, recommendations, and similarity matching. It is available through the Gemini API and Vertex AI, supports up to 2,048 input tokens and between 128 and 3,072 output dimensions, and is scheduled to shut down on May 14, 2028.

What is Gemini Embedding?

Gemini Embedding is a text embedding model from Google DeepMind. Its canonical model identifier is gemini-embedding-001, and it is available through the Gemini API and Google Cloud Vertex AI. Instead of generating a conversational answer, the model converts text into a numerical vector, also called an embedding.

An embedding represents important patterns in the meaning and use of the input text. A search application can compare the vector for a user’s question with vectors created from documents or passages. Texts with related meanings should generally be located closer together in vector space than unrelated texts, even when they do not use the same keywords.

This makes Gemini Embedding useful as one component in a larger application. It can help a document system find relevant passages, provide candidate results for a retrieval-augmented generation (RAG) pipeline, group similar content, detect near-duplicates, classify text, or recommend related items. It does not independently write the final answer, summarize a document, browse the web, or operate tools.

Where Gemini Embedding fits in Google’s lineup

Gemini Embedding belongs to the representation and retrieval part of Google’s Gemini catalog rather than the generative model lineup. Generative Gemini models accept prompts and return text or other generated content. Gemini Embedding accepts text and returns vectors intended for downstream software.

That distinction matters when evaluating the model. It is not a smaller chat model and should not be selected for an assistant that needs to explain information directly. Its purpose is to make text computationally comparable. A typical architecture might use Gemini Embedding to retrieve relevant passages, then pass those passages to a separate generative model that produces an answer.

Google’s documentation describes Gemini Embedding 2 as the recommended successor for applications that need a newer embedding option, while gemini-embedding-001 is scheduled to shut down on May 14, 2028. The shutdown date makes migration planning part of the model’s practical positioning, particularly for new systems expected to operate for several years.

Technical specifications and limits

SpecificationVerified detail
Model identifiergemini-embedding-001
ProviderGoogle DeepMind
Primary inputText
Primary outputNumerical text embedding vectors
Maximum input length2,048 tokens
Configurable output dimensions128 to 3,072
Recommended dimensions768, 1,536, or 3,072
AccessGemini API and Vertex AI
Scheduled shutdownMay 14, 2028

The 2,048-token input limit applies to each text input and affects how long a document passage can be embedded in one request. Longer documents generally need to be divided into smaller chunks before indexing. Chunking strategy can affect retrieval quality: chunks that are too short may lose context, while chunks that are too long may contain several unrelated topics and make retrieval less precise.

The output dimension controls how many numbers each vector contains. Gemini Embedding can produce vectors from 128 through 3,072 dimensions. Google recommends 768, 1,536, or 3,072 dimensions for common deployments. Larger vectors can preserve more representational information, but they require more storage and may increase the resource requirements of vector search. Smaller vectors can reduce those costs, although each application should test the quality trade-off on its own data.

Why configurable dimensions matter

Gemini Embedding uses Matryoshka Representation Learning, which allows applications to request smaller vector representations by truncating the full representation. In practical terms, a team can choose a vector size that fits its database and latency budget instead of being locked to one fixed dimensionality.

This flexibility is useful when an application has millions of indexed passages. Reducing the number of dimensions can lower storage requirements and reduce the amount of data processed during similarity searches. However, changing dimensions is not a cosmetic setting. The vectors stored in the index and the vectors generated for incoming queries must use a compatible dimensionality and configuration.

Teams should also keep their embedding model, task configuration, dimensionality, and normalization approach consistent between indexing and querying. If documents are indexed with one setup and user queries are embedded with another, similarity scores may no longer be directly comparable. The model therefore offers deployment flexibility, but that flexibility needs to be managed as part of the vector database design.

What Gemini Embedding is good for

Keyword search depends heavily on exact word overlap. Embedding-based search instead compares vector representations, allowing a query and a document to match when they express a similar idea with different wording. For example, a support search system may retrieve a document about “resetting account credentials” for a user who asks how to “change a forgotten password.”

The embedding model does not replace the entire search system. An application still needs an index, a similarity-search method, ranking logic, and a way to display or use the retrieved results. Gemini Embedding supplies the vector representation that makes semantic comparison possible.

RAG and document grounding

In a RAG system, documents are split into passages and embedded in advance. When a user asks a question, the question is embedded and compared with the stored passage vectors. The most relevant passages can then be supplied to a separate generative model as context.

This approach can help a generative system use a private document collection without placing the entire collection into every prompt. Gemini Embedding is suitable for the retrieval stage, but it does not generate the grounded response itself. Google also documents embedding-based retrieval through services such as File Search and Vertex AI data services.

Classification, clustering, and recommendations

Embedding vectors can serve as input features for conventional or custom application logic. A team can group similar documents, identify duplicate or related records, classify support tickets, find comparable products, or recommend content based on similarity. These applications may require additional models, rules, or labeled data; Gemini Embedding provides the representation rather than a complete classification or recommendation product.

Input, output, reasoning, and tool support

Gemini Embedding is text-only. According to the supplied model specifications, it does not accept images, audio, video, or PDFs as native embedding inputs. A PDF must therefore be converted into suitable text before this model can process its contents. The model also does not produce text, images, audio, video, or other direct media output. Its output is an embedding vector.

It is not a reasoning model in the conversational sense. It does not work through a problem and return an explanation, and it has no documented maximum output-token setting because it does not generate tokens. Similarly, it does not provide native code generation, function calling, web search, or tool use. A developer can use its vectors inside a sophisticated reasoning or coding application, but those capabilities would come from other components.

The supplied catalog rates its reasoning and coding suitability at a low editorial level, not as provider-published benchmark results. The same catalog gives it high relative speed and cost scores because embedding requests return compact vector data and are designed for retrieval workloads. These scores are evaluations for comparison, not guarantees of latency, throughput, or total system cost.

Pricing and API access

Gemini Embedding is exposed through the Gemini API and Vertex AI. The supplied Vertex AI pricing documentation lists online requests at $0.00015 per 1,000 input tokens and batch requests at $0.00012 per 1,000 input tokens. Embedding output is not charged separately under the cited Vertex AI pricing information.

Batch processing is the lower-cost option in the supplied pricing record and is appropriate for indexing a large document collection when immediate results are not required. Online requests are more appropriate for interactive query embedding, where an application needs to process a user’s search immediately. Actual costs also depend on how often documents are re-embedded, how documents are chunked, and which billing surface and region are used.

The published rates in this comparison are Vertex AI rates. Gemini API availability and pricing may differ from Vertex AI, so a project should verify the current terms for its chosen access route before estimating production costs.

Important limitations and migration concerns

The most immediate technical limitation is the 2,048-token input ceiling. Applications that process long reports, books, transcripts, or large code repositories must split them into passages. Splitting content creates an application-design problem: the system must preserve enough local context for useful retrieval without making every passage unnecessarily large.

The model is also limited to text. It is not the right choice for a single embedding space covering images, audio, video, and text. Applications with genuine multimodal retrieval requirements should evaluate a model designed for those modalities, including Google’s documented successor direction where appropriate.

Finally, gemini-embedding-001 has a scheduled shutdown date of May 14, 2028. A system launched today should keep the embedding pipeline replaceable: retain source documents, record the model and dimensionality used for each index, and plan how to regenerate vectors. Migration may require rebuilding the vector index because embeddings from a successor model should not automatically be assumed to be numerically compatible with the existing index.

When to choose Gemini Embedding

Choose Gemini Embedding when the primary requirement is text representation for semantic search, RAG retrieval, document similarity, clustering, classification features, or recommendations. It is especially practical when configurable dimensions are useful for balancing retrieval quality, vector storage, and search cost, and when the 2,048-token input limit fits the application’s chunking strategy.

It may be a good fit for a text-only retrieval pipeline that already has a separate generative model for answering questions. Its documented Vertex AI pricing also makes batch indexing relatively inexpensive compared with interactive processing, although the complete system cost includes storage, vector search, data processing, and any generation model used after retrieval.

Choose another option when the model must directly answer questions, generate code, call tools, browse the web, process native image or audio inputs, or avoid a known shutdown date. A newer embedding model may also be more appropriate for a new long-lived deployment if it offers the required modality coverage or a clearer lifecycle. Regardless of the alternative, the comparison should be made using the application’s own retrieval tests rather than assuming that a larger vector or newer model will automatically produce better results.


Answers to Frequently Asked Questions

When will gemini-embedding-001 shut down, and what should users do?
gemini-embedding-001 is scheduled to shut down on May 14, 2028. Teams should retain their source documents, record the model and dimensionality used for each index, and plan to regenerate embeddings and potentially rebuild the vector index when migrating to a successor such as Gemini Embedding 2.
What vector dimensions does Gemini Embedding support?
Gemini Embedding supports configurable output dimensions from 128 to 3,072. Google recommends 768, 1,536, or 3,072 dimensions. The same model, task configuration, dimensionality, and normalization approach should be used when indexing documents and embedding queries.
What is the model identifier and input limit for Gemini Embedding?
The canonical model identifier is gemini-embedding-001. It accepts text inputs of up to 2,048 tokens per request, so longer documents usually need to be split into smaller passages before embedding.
Can Gemini Embedding generate answers or process images and PDFs?
No. Gemini Embedding is a text-only representation model that outputs embedding vectors. It does not generate conversational answers, browse the web, call tools, or natively process images, audio, video, or PDFs. PDF content must be converted to text first.
What is Gemini Embedding used for?
Gemini Embedding converts text into numerical vectors for semantic search, RAG retrieval, document similarity, clustering, classification, duplicate detection, and recommendations. It provides the representation for these applications but does not generate final answers or operate tools.


Sources 6
Provider

About Google DeepMind