What text-embedding-ada-002 does
text-embedding-ada-002 is an OpenAI model for turning text into an embedding: a fixed-length list of numbers that represents patterns of meaning in the input. Applications can compare these vectors to estimate how closely two pieces of text are related.
For example, a search system can embed a user's question and compare it with embeddings stored for documents. A document discussing how to reset a password may be retrieved for a query such as “I cannot access my account,” even when the words do not match exactly. The model itself does not write an answer; it supplies the numerical representation that another application can use for matching, ranking, or classification.
text-embedding-ada-002 is accessed through OpenAI's embeddings endpoint rather than a text-generation endpoint. Its primary role is therefore semantic representation, not conversation, reasoning, code generation, or content creation.
Verified specifications
| Specification | text-embedding-ada-002 |
|---|---|
| Provider | OpenAI |
| Model type | Text embedding |
| Release date | December 15, 2022 |
| Input | Text |
| Maximum input | 8,192 tokens |
| Output | Numerical embedding vector |
| Vector size | 1,536 dimensions |
| Listed price | $0.10 per 1 million input tokens |
| Text generation | Not supported |
| Image, audio, and video output | Not supported |
The vector has the same 1,536 dimensions whether the input is short or long. The input must still remain within the model's 8,192-token limit. A token is a unit used to process text and may represent a whole word, part of a word, punctuation, or a space-related fragment, so the token count is not always identical to the number of words.
Where it fits in OpenAI's current lineup
OpenAI identifies text-embedding-ada-002 as an older embedding model. It remains listed and accessible through the embeddings API, but it is no longer the default starting point for many new embedding systems. The newer text-embedding-3-small and text-embedding-3-large models are its principal successors.
OpenAI's published information states that text-embedding-ada-002 was not being deprecated when the text-embedding-3 family was introduced. That does not make its embeddings interchangeable with newer vectors, however. Each model creates its own embedding space, so an index built with ada-002 normally cannot be queried directly with vectors produced by a different model.
In practical terms, ada-002 occupies a compatibility-oriented position. It can remain useful for an existing application whose database, ranking behavior, and evaluation process were built around its 1,536-dimensional vectors. For a new project, the newer family may offer a better balance of retrieval quality, price, and configuration options.
Pricing and cost trade-offs
The supplied OpenAI model information lists text-embedding-ada-002 at $0.10 per 1 million input tokens. This is an input-only embedding price: the model returns vectors rather than generated text, so there is no conventional output-token charge in the cited pricing information.
Price alone does not determine the cost of an embedding project. A system also has to process the initial document collection, embed future documents, and embed incoming user queries. Storage and vector-search infrastructure are separate operational costs. If an existing application already contains a large ada-002 index, migrating may require re-embedding both the stored documents and future queries, which can make compatibility more valuable than a lower per-token price.
For new deployments, OpenAI describes text-embedding-3-small as less expensive and stronger on published comparisons than ada-002. That comparison favors the newer model when the project has no legacy index to preserve. The exact business choice should still be tested on the application's own documents and queries rather than inferred from a general benchmark alone.
Main use cases
- Semantic search: Find documents, help-center articles, or records by meaning rather than exact wording.
- Retrieval-augmented generation: Retrieve relevant passages for a separate text-generation model before constructing an answer.
- Clustering: Group documents, tickets, or other text according to semantic similarity.
- Recommendations: Match users, products, documents, or topics based on related text descriptions.
- Anomaly detection: Identify items whose vectors are unusually distant from the normal pattern.
- Classification: Supply embeddings to a downstream classifier that assigns labels such as topic, department, or priority.
A typical retrieval workflow splits source documents into manageable passages, sends those passages to the embeddings endpoint, stores the resulting vectors in a vector database, and later embeds each incoming query with the same model. A similarity measure such as cosine similarity can then rank passages for the query.
Strengths and limitations
Strengths
- Established compatibility: Many existing search and retrieval systems may already be designed around its 1,536-dimensional vectors.
- Useful input capacity: The 8,192-token limit is substantially more accommodating than the earlier 2,048-token embedding limit that OpenAI described when introducing the model.
- Broad applicability: The same embedding approach can support search, retrieval, clustering, recommendations, anomaly detection, and classification.
- Simple output format: Applications receive a fixed-size numerical vector that can be stored and compared using standard vector-search infrastructure.
Limitations
- Text only: The model accepts text and returns numerical vectors. It does not natively produce text, images, audio, or video.
- No generation or reasoning endpoint: It should not be selected when the application needs answers, explanations, code, tool use, or conversational interaction from the model itself.
- Fixed dimensions: Unlike newer text-embedding-3 models, ada-002 does not provide a configurable dimensions parameter in the supplied specifications.
- Migration friction: Vectors from ada-002 are not interchangeable with vectors from other embedding families. Moving to another model generally requires re-embedding stored content and incoming queries.
- Older efficiency position: OpenAI's newer embedding models are positioned as stronger or more efficient choices for many new projects.
Input, output, and capability boundaries
text-embedding-ada-002 supports text input only. Its output is an embedding vector, not a response that a user can read directly. The 1,536 numbers encode useful relationships for downstream software, but they are not a generated summary, translation, classification label, or explanation.
The model has no cited support for multimodal input, image understanding, audio processing, video processing, tool calling, streaming responses, structured JSON output, or text generation. These are not merely missing user-interface features; they are outside the model's documented role as an embedding model.
Its reasoning and coding capabilities should therefore be considered not applicable. It can represent text about a programming problem or reasoning task for retrieval or classification, but it does not solve the problem or execute code. If an application needs a generated answer after retrieval, ada-002 would normally be one component of a larger pipeline rather than the answering model.
When to choose text-embedding-ada-002
Choose text-embedding-ada-002 when an existing production system already depends on its vectors and the priority is avoiding a disruptive index migration. It can also be a reasonable choice when a legacy application has been evaluated with ada-002 and its retrieval quality, storage schema, and ranking thresholds are already stable.
Before retaining it, confirm that the model is still available for the account and deployment path, that the 8,192-token limit is sufficient for the selected text chunks, and that the $0.10-per-million-input-token price fits the current budget. It is especially important to embed both documents and queries with the same model.
For a new search, recommendation, or retrieval system, compare ada-002 with text-embedding-3-small and text-embedding-3-large on representative data. The newer models are more appropriate when the project prioritizes current retrieval quality, lower pricing, configurable dimensions, or a new vector index without legacy compatibility requirements. A large or complex migration should be justified by measured improvements in search relevance, operational cost, storage, or maintenance.
Migration considerations
Migration is not a matter of changing the model name for future queries while keeping the old vectors. A consistent replacement workflow normally selects the new model, re-embeds the stored source content, rebuilds or updates the vector index, and changes query embedding at the same time. Mixing ada-002 vectors with vectors from another embedding family can invalidate similarity comparisons.
Evaluate the replacement using real searches and labeled relevance judgments where possible. Check whether the vector database accepts the new dimensions, whether storage and indexing costs change, and whether application thresholds need adjustment. Keep the old index available during testing if rollback or result comparison is important.
Bottom line
text-embedding-ada-002 remains a usable OpenAI embedding model for text-to-vector workflows, with 1,536-dimensional outputs, an 8,192-token input limit, and a listed price of $0.10 per 1 million input tokens. Its strongest practical justification today is compatibility with established systems. For new projects, the newer text-embedding-3 models are generally the more appropriate alternatives, while ada-002 remains relevant when preserving an existing vector index and retrieval pipeline is more important than moving to a newer embedding space.

