text-embedding-3

text-embedding-3-large

by OpenAI · Current and available through the OpenAI API

OpenAI's text-embedding-3-large converts text and code into semantic vectors for search, RAG, recommendations, clustering, classification and similarity detection. It supports up to 8,192 input tokens, returns up to 3,072 dimensions, allows shorter vectors through the dimensions parameter and costs $0.13 per 1 million input tokens.

Embeddings Reasoning Coding
text-embedding-3-large is a specialized OpenAI model for turning text or code into numerical representations of meaning. Rather than generating an answer like a conversational model, it helps software find related content, compare documents, organize information and retrieve relevant passages. It accepts inputs of up to 8,192 tokens, produces vectors with a default length of 3,072 dimensions, and is priced at $0.13 per 1 million input tokens.
Outputs

What text-embedding-3-large can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

1/10 Reasoning
2/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family text-embedding-3
Model type Embedding
Context window 8K tokens
Release date 2024-01-25
Status Current and available through the OpenAI API
Model notes

OpenAI released the model on January 25, 2024. It is a specialized embedding model rather than a generative language model. The default embedding length is 3,072 dimensions, but the dimensions parameter can request shorter vectors, including configurations such as 1,024 or 256 dimensions. OpenAI reports 54.9% average MIRACL performance and 64.6% average MTEB performance at launch. The model accepts text and returns numerical vectors through the Embeddings API. The current model page lists Batch API availability. Editorial scores are comparative estimates for this specialized model category, not provider benchmarks.

Cost

Model pricing

Input $0.13 per 1 million input tokens
Output No separate output-token charge; the model returns embedding vectors
Model guide

text-embedding-3-large: Features, Pricing, Dimensions and API Use

text-embedding-3-large is OpenAI's high-capacity text embedding model. It converts text and code into numerical vectors for semantic search, retrieval-augmented generation, recommendations, clustering, classification, similarity matching and anomaly detection. It supports up to 8,192 input tokens, returns vectors of up to 3,072 dimensions, and can produce shorter vectors through the dimensions parameter.

What is text-embedding-3-large?

text-embedding-3-large is OpenAI's high-capacity third-generation text embedding model. An embedding is a list of numbers that represents the semantic content of text or code. Texts with similar meanings generally produce vectors that are close together when compared with an appropriate similarity measure.

This makes the model useful for applications that need to find or compare information rather than write prose. For example, a support system can embed a customer's question and compare it with embeddings of help-center articles. A retrieval-augmented generation (RAG) system can use the same process to locate relevant passages before passing them to a separate language model for an answer.

The model is available through OpenAI's Embeddings API and is intended for developer applications. It was released on January 25, 2024 and remains listed as a current model in OpenAI's API catalog.

What the model does and does not do

text-embedding-3-large accepts text input and returns a numerical embedding vector. It does not return conversational text, images, audio or video. It is therefore better understood as a component inside a search, recommendation or machine-learning pipeline than as a chatbot.

SpecificationVerified detail
ProviderOpenAI
Model typeText embedding
InputText, including text used to represent code
OutputNumerical embedding vector
Maximum input8,192 tokens
Default vector length3,072 dimensions
API endpoint/v1/embeddings
Separate text outputNo

The 8,192-token limit applies to the input sent for embedding. It is not a conversational context window, and there is no maximum output-token setting because the output is a vector rather than generated language.

Adjustable vector dimensions

A major practical feature is the ability to request shorter vectors with the dimensions parameter. The default output contains 3,072 numbers, but OpenAI has documented configurations such as 1,024 and 256 dimensions.

Shorter vectors can reduce storage requirements, memory consumption and the cost of operating a vector database. They can also make similarity searches more efficient when an application handles a large document collection. The trade-off is that reducing the vector length may reduce semantic accuracy for some workloads.

There is no universally correct dimension setting. A product catalog, multilingual knowledge base and code-search system may respond differently to compression. Teams should evaluate candidate dimensions against representative queries and their actual success criteria, such as relevant-document recall or classification accuracy. The dimensions setting is an engineering trade-off, not an automatic quality improvement.

Performance and quality

OpenAI reported average MIRACL performance of 54.9% and average MTEB performance of 64.6% when the model launched. These are provider-reported launch benchmarks, not guarantees for every dataset or application. They positioned text-embedding-3-large as OpenAI's strongest embedding option at that time.

Benchmark results should be treated as one comparison point. Real-world retrieval quality also depends on how documents are divided into chunks, how queries are written, which distance metric is used, whether vectors are normalized, how metadata filters are applied and whether a reranking step is added. A model with strong benchmark results can still produce poor search results if the source documents are badly segmented or the retrieval pipeline is not evaluated.

For this reason, the model's practical advantage is best assessed with an application-specific test set. Measure whether the retrieved passages actually answer representative user questions, rather than relying only on a general benchmark score.

Pricing and operating costs

The listed price is $0.13 per 1 million input tokens. Embedding usage is charged according to input tokens. There is no separate output-token charge because the model returns vectors rather than generated prose.

The API price is only one part of the total cost. A system using the default 3,072-dimensional output must also store and index those vectors. At large scale, vector-database storage, indexing, backups and search infrastructure can become important. Requesting shorter vectors may reduce these operational costs, although the result should be checked for any loss in retrieval quality.

OpenAI also lists the model as available through the Batch API. Batch processing can be useful when a large document collection needs to be embedded asynchronously instead of immediately during an interactive request. The research supplied here does not specify a separate Batch API price, so the standard listed input price should not be interpreted as a complete estimate of infrastructure cost.

Supported capabilities and limitations

text-embedding-3-large supports semantic representation of text and code, but it does not provide the capabilities associated with generative language models.

  • Text input: Supported.
  • Code input: Supported as text for representation, search or classification.
  • Text output: Not supported; the model returns vectors.
  • Image, audio and video input: Not supported.
  • Image, audio and video output: Not supported.
  • Reasoning: Not a reasoning model. It does not analyze a question and produce a chain of conclusions.
  • Conversational generation: Not supported.
  • Tool or function calling: Not supported as a native model capability.
  • Web search: Not built in.
  • Streaming: Not listed as supported for this model.

These limitations do not prevent an application from combining embeddings with other services. For example, a system can use text-embedding-3-large for retrieval and then send the selected passages to a separate generative model. In that design, the embedding model handles similarity and the other model handles explanation or response generation.

Best use cases

The model is a good fit when an application needs a high-quality semantic representation rather than a written response.

  • Semantic search: Find documents based on meaning even when the query and document use different words.
  • Retrieval-augmented generation: Retrieve relevant source passages before a separate language model generates an answer.
  • Multilingual retrieval: Compare documents and queries across supported languages.
  • Recommendations: Identify related articles, products or documents by comparing their vectors.
  • Duplicate and similarity detection: Find records that express similar ideas despite wording differences.
  • Clustering: Group documents or user-generated content by semantic similarity.
  • Classification: Use vectors as features for categorization systems.
  • Anomaly detection: Identify items that are unusually distant from the normal semantic pattern.
  • Code search: Represent code or code-related text for similarity-based discovery, provided the surrounding application handles language-specific indexing and evaluation.

When to choose text-embedding-3-large

Choose text-embedding-3-large when retrieval quality is more important than using the smallest possible vector, and when the application benefits from a high-capacity embedding with configurable output size. It is particularly suitable for multilingual search, large knowledge bases, RAG pipelines and similarity systems where the default 3,072-dimensional representation can justify its storage and indexing cost.

The model is also a sensible choice when an application needs flexibility. Developers can begin with the default dimensions and later test shorter vectors without changing to a completely different model family. That makes it possible to balance quality, latency and infrastructure cost using the same API model.

A smaller or less capable embedding option may be more appropriate when storage and search cost dominate the design, when the corpus is small and simple, or when testing shows that a shorter vector provides adequate recall. Conversely, a conversational or reasoning model is more appropriate when the system must explain results, follow instructions, call tools or generate a final answer. text-embedding-3-large should normally supply the retrieval layer rather than replace those models.

Practical design considerations

Embedding quality depends on the text supplied to the model. Before sending documents, decide how large each chunk should be, whether headings and metadata should be included, and how to preserve enough surrounding context for a retrieved passage to be useful. The 8,192-token input limit is a maximum for an individual embedding request, not a recommendation to place an entire large document into one vector.

Search quality should also be evaluated at the complete pipeline level. Compare the retrieved results for real user queries, test metadata filtering, examine failures involving synonyms or multilingual wording, and determine whether reranking is needed. If shorter vectors are used, compare them directly with the 3,072-dimensional default on the same evaluation set.

Because the model returns no explanation with its vector, applications should retain document identifiers and source metadata alongside each embedding. That allows the system to show where a result came from and lets a separate response model use the retrieved content safely.

Bottom line

text-embedding-3-large is a specialized OpenAI model for converting text and code into searchable numerical representations. Its main differentiators are the 3,072-dimension default output, support for up to 8,192 input tokens, adjustable vector dimensions and a listed price of $0.13 per 1 million input tokens. It is strongest as the semantic retrieval or similarity layer in an application, not as a standalone conversational system.

Its main trade-off is capacity versus operating cost: the default vectors can support high-quality retrieval but require more storage and indexing resources than shorter representations. Teams should select the dimensions setting, chunking strategy and surrounding retrieval pipeline through application-specific evaluation.


Answers to Frequently Asked Questions

Can text-embedding-3-large generate answers or support conversational features?
No. text-embedding-3-large only produces numerical embedding vectors and does not generate text, perform web searches, call tools, process images, audio or video, or provide conversational reasoning. It can supply retrieved content to a separate generative model that produces the final answer.
How do I use text-embedding-3-large through the API?
Use OpenAI's Embeddings API at the /v1/embeddings endpoint, provide text input and optionally specify the dimensions parameter to request a shorter vector. The response contains a numerical embedding that can be stored and compared in a vector database.
How much does text-embedding-3-large cost?
The listed price is $0.13 per 1 million input tokens. There is no separate output-token charge because the model returns an embedding vector rather than generated text. Storage, indexing, backups and vector-database search infrastructure add to the overall operating cost.
What is text-embedding-3-large used for?
text-embedding-3-large converts text or code into numerical vectors that represent semantic meaning. It is used for semantic search, retrieval-augmented generation, recommendations, similarity detection, clustering, classification, anomaly detection and code search.
What are the dimensions and input limit of text-embedding-3-large?
The model returns 3,072-dimensional vectors by default and accepts up to 8,192 input tokens per request. Developers can request shorter vectors, including documented configurations such as 1,024 or 256 dimensions, using the dimensions parameter.


Sources 3
Provider

About OpenAI