What is text-embedding-3-small?
text-embedding-3-small is an OpenAI model that converts text into an embedding: a list of numbers representing the semantic characteristics of the input. Texts with related meanings tend to produce vectors that are closer together when measured with a suitable similarity method. This lets software find relevant content without requiring an exact keyword match.
For example, a search for “how do I reset my password?” can be matched with a support document titled “Recovering access to your account,” even though the two texts use different words. The model itself does not write the answer. It produces the vector representation that another application can use to retrieve relevant information, classify content, or calculate similarity.
OpenAI released text-embedding-3-small on January 25, 2024, as part of the text-embedding-3 model family. It is listed in OpenAI’s current model catalog and is accessed through the Embeddings API using the model identifier text-embedding-3-small.
Capabilities and specifications
| Specification | Details |
|---|---|
| Provider | OpenAI |
| Model type | Text embedding |
| Release date | January 25, 2024 |
| Input | Text |
| Output | Embedding vector |
| Default vector size | 1,536 dimensions |
| Maximum input | 8,192 tokens |
| Knowledge cutoff | September 2021 |
| Endpoint | OpenAI Embeddings API |
| Canonical model ID | text-embedding-3-small |
The default output contains 1,536 numerical values. OpenAI also supports reducing the vector length with the dimensions parameter. Shorter vectors require less storage and can reduce vector-database costs, but reducing dimensions may involve a quality trade-off for a particular search or classification workload. The appropriate size should be tested against the application’s own evaluation data.
The 8,192-token input limit applies to an individual embedding request input. Long documents therefore usually need to be divided into smaller passages before embedding. The chunking strategy affects retrieval quality: chunks that are too short may lose context, while chunks that are too long may combine several unrelated topics.
What it can and cannot do
text-embedding-3-small accepts text and returns vectors. It does not produce natural-language responses, images, audio, or video. It is not a conversational model and does not provide reasoning, web search, tool calling, or function execution. It also does not directly answer a user’s question or generate a completed summary.
Its output is useful because applications can compare vectors. A search system may embed a user query, compare it with previously stored document vectors, and return the closest passages. A retrieval-augmented generation system can then pass those passages to a separate text-generation model that writes the final answer. In this arrangement, text-embedding-3-small handles representation and retrieval rather than generation.
The model has no separate generated-text output limit because it does not generate prose. The important size constraints are the maximum input length and the number of dimensions in each returned vector.
Performance and positioning
text-embedding-3-small is positioned as a cost-efficient embedding option for high-volume workloads. It is intended for applications that need many vector representations without paying the higher cost or storage overhead that can accompany larger embedding configurations.
OpenAI reports that text-embedding-3-small improves on text-embedding-ada-002 in selected retrieval evaluations. In the provider’s published comparison, it scored 44.0% on the MIRACL multilingual retrieval benchmark and 62.3% on the MTEB English-task benchmark, compared with 31.4% and 61.0%, respectively, for text-embedding-ada-002. These are provider-reported benchmark results, not a guarantee of performance on every dataset. Results can vary with language, document quality, chunking, similarity metric, and application design.
OpenAI also offers the larger text-embedding-3-large. The larger model may be more appropriate when maximum retrieval quality is more important than minimizing price, storage, or processing requirements. text-embedding-3-small is generally the more practical starting point for large collections, latency-sensitive searches, and applications whose quality requirements can be met with a smaller model.
Pricing
The standard listed price is $0.02 per 1 million input tokens. Billing is based on the text supplied for embedding. There is no separate charge for generated output because the response contains vectors rather than generated text.
At this price, the model is suitable for indexing substantial quantities of documents and for applications that embed frequent user queries. The total system cost can still include storage, vector-database operations, application hosting, and any separate language model used to generate answers. If an application uses retrieval-augmented generation, the embedding charge is only one part of the overall workflow cost.
OpenAI’s model information also lists Batch API availability. Batch processing can be useful for large indexing jobs where requests do not need to be handled interactively, although the supplied model information does not specify a separate batch price.
What is text-embedding-3-small used for?
- Semantic search: Find documents, help-center articles, products, or records by meaning rather than exact wording.
- Retrieval-augmented generation: Retrieve relevant passages for a separate generative model before it formulates an answer.
- Clustering: Group documents, tickets, reviews, or other text according to similarity.
- Recommendations: Match users, products, articles, or other items using vector similarity.
- Anomaly detection: Identify inputs that are unusually different from an established collection.
- Duplicate detection: Locate documents or messages that express similar content even when they are not identical.
- Classification: Use vector similarity or a separate classifier to assign categories to text.
- Code and content similarity: Compare text or code samples when a semantic representation is more useful than literal string matching.
A typical document-search workflow stores an embedding for each document passage in a vector database. When a user searches, the application embeds the query with the same model, compares the query vector with stored vectors, and returns the nearest matches. Using the same embedding model for indexing and querying is important because vectors from different models are not generally interchangeable.
Limitations and engineering trade-offs
The model does not remove the need for application design. Retrieval quality depends on how documents are cleaned and divided, which similarity metric is used, whether vectors are normalized, how many results are retrieved, and how relevance thresholds are chosen. A technically correct embedding model can still produce poor search results if the indexed content is incomplete or the chunks lack useful context.
The model’s documented knowledge cutoff is September 2021. This does not prevent it from embedding newer text supplied by an application, but it means the underlying model does not have learned knowledge of events after that date. Applications that index current information can still retrieve that information because the current documents themselves are provided as input.
Changing embedding models normally requires re-embedding the collection. Existing vectors from another model should not simply be mixed with new vectors, and similarity thresholds may need to be retuned after migration. This can make model replacement an operational project for large indexes.
Shortening the vector with dimensions can lower storage and infrastructure costs, but it may reduce retrieval accuracy. The trade-off is application-specific, so teams should compare recall, ranking quality, latency, and storage cost before selecting a shortened representation.
Reasoning, coding, and tool support
text-embedding-3-small has no conversational reasoning capability in the usual sense. It does not plan a multi-step answer, interpret instructions as a chat assistant, or produce a natural-language explanation. Its role is to encode input text into a numerical representation.
It can be used with code or technical documents for similarity and retrieval, but it is not a code-generation model and should not be selected for writing, debugging, or executing programs. It also has no native tool or function-calling support, web-search capability, streaming output, or structured text output. Those functions must be provided by the surrounding application or by another model.
When to choose text-embedding-3-small
Choose text-embedding-3-small when the primary requirement is affordable, large-scale text representation. It is a strong candidate for semantic search, document retrieval, support portals, content organization, recommendation features, and other systems where every query or document must be converted into a vector.
- Choose it for a cost-sensitive index containing many documents or frequent searches.
- Choose it when 1,536 default dimensions provide sufficient quality and storage is an important consideration.
- Choose it when you need configurable vector length and can evaluate the effect of shortening dimensions.
- Choose it when the application needs multilingual or English retrieval and can validate performance on its own content.
A larger embedding model may be more appropriate when retrieval quality is the overriding priority and additional cost or storage is acceptable. A generative model is more appropriate when the application must answer questions, summarize documents, write code, or follow conversational instructions. For image, audio, or video inputs, a model with those modalities is required instead.
Availability and current status
text-embedding-3-small is currently listed in OpenAI’s model catalog. The available alias is text-embedding-3-small, and the supplied documentation does not identify a separate dated snapshot for ordinary use. It is accessed through OpenAI’s v1/embeddings endpoint.
For developers, the main decision is therefore not whether the model can generate an answer, but whether its price, vector size, input limit, and retrieval quality fit the application. When those requirements align, it provides a compact foundation for search and similarity systems; when they do not, a larger embedding model or a generative model may be a better fit.

