Qwen3-Embedding

text-embedding-v3

by Qwen · Available; retained primarily for compatibility with existing v3 embedding indexes

Alibaba Cloud text-embedding-v3 converts multilingual text into vectors for semantic search, RAG, recommendations, clustering, and classification. It supports inputs up to 8,192 tokens, more than 50 major languages, and 1,024-, 768-, or 512-dimensional embeddings. Its key role is maintaining compatibility with existing v3 indexes, while Alibaba Cloud recommends text-embedding-v4 for many new deployments. The review covers pricing, rate limits, supported modalities, limitations, and selection guidance.

Text Embeddings
text-embedding-v3 is a general-purpose text embedding model from Alibaba Cloud Model Studio. Instead of generating an answer in natural language, it converts supplied text into numerical vectors that applications can compare for semantic similarity. That makes it useful for search, retrieval-augmented generation (RAG), recommendations, clustering, and classification. The model supports input of up to 8,192 tokens, more than 50 major languages, and configurable embedding dimensions of 1,024, 768, or 512. Its strongest practical use case is maintaining compatibility with existing v3 indexes; Alibaba Cloud currently recommends text-embedding-v4 for many new text-search and RAG deployments.
Outputs

What text-embedding-v3 can produce

Text Embeddings
Inputs

What it can understand

Text
Model profile

Performance characteristics

8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3-Embedding
Model type Embedding
Context window 8K tokens
Status Available; retained primarily for compatibility with existing v3 embedding indexes
Knowledge cutoff notes

Alibaba Cloud does not publish a separate knowledge-cutoff date for this embedding model. Embedding models are intended to transform supplied input rather than answer questions from a documented training cutoff.

Model notes

General-purpose multilingual text embedding model developed by Tongyi Lab and served through Alibaba Cloud Model Studio. It accepts text and returns vector embeddings. Supported dimensions are 1,024 by default, 768, and 512. The synchronous API supports input text up to 8,192 tokens and more than 50 major languages. Alibaba Cloud currently recommends text-embedding-v4 for new text-search and RAG deployments, while text-embedding-v3 remains useful for maintaining dimension compatibility with existing indexes. The model page marks structured outputs, function calling, web search, context caching, batch inference, and fine-tuning as unsupported. Editorial speed and cost scores are comparative estimates for embedding workloads, not vendor ratings.

Cost

Model pricing

Input $0.07 per 1 million input tokens in the Singapore international deployment
Output Free
Model guide

text-embedding-v3: Multilingual Embeddings for Search and Existing Vector Indexes

Alibaba Cloud's text-embedding-v3 converts text into multilingual vector embeddings for semantic search, retrieval-augmented generation, recommendations, clustering, and classification. Its main practical advantage is compatibility with existing v3 vector indexes, while text-embedding-v4 is the provider's recommended choice for many new deployments.

What is text-embedding-v3?

text-embedding-v3 is a text embedding model provided through Alibaba Cloud Model Studio and developed by Tongyi Lab. It transforms text into a vector: a list of numbers that represents aspects of the text's meaning. Software can compare those vectors to find passages, documents, queries, or user interests that are semantically similar, even when they do not use exactly the same words.

This is different from a conversational language model. text-embedding-v3 does not primarily write explanations, summarize documents, or answer questions. An application typically sends text to the model, stores the resulting vector in a vector database, and later compares a query vector with stored vectors. In a RAG system, for example, the closest matching passages can be retrieved and passed to a separate generative model.

The model belongs to Alibaba Cloud's Qwen3-Embedding family. In the provider's current positioning, text-embedding-v3 remains useful for existing deployments and indexes, while text-embedding-v4 is recommended for new text-search and RAG deployments where a migration is practical.

Core specifications and supported modalities

Specificationtext-embedding-v3
ProviderAlibaba Cloud Model Studio
Model typeText embedding
InputText
OutputVector embedding
Maximum synchronous input8,192 tokens
Supported languagesMore than 50 major languages
Supported dimensions1,024 by default; 768 or 512 also available
Image, audio, and video inputNot supported
Generated text, image, audio, or video outputNot supported

The 8,192-token limit applies to the input accepted by the synchronous embedding API. A token is a unit used to process text and may be a whole word, part of a word, punctuation, or another language-specific unit. Long documents therefore need to be split into smaller passages before embedding. Splitting is also useful for retrieval because search systems generally work better when individual vectors represent focused sections rather than entire books or large collections of unrelated material.

Embedding dimensions affect storage and similarity-search costs. The default 1,024-dimensional representation may preserve more detail, while 768 and 512 dimensions reduce the size of stored vectors. The appropriate choice depends on the index design and the compatibility requirements of the application. Changing dimensions for an existing index generally requires an indexing plan because vectors with different dimensions cannot simply be mixed in the same comparison workflow.

What text-embedding-v3 is good for

  • Semantic search: Match a user's query with documents based on meaning rather than exact keyword overlap.
  • Retrieval-augmented generation: Retrieve relevant passages before a separate language model generates an answer.
  • Recommendations: Represent content, products, or user interests as vectors and compare their relationships.
  • Clustering: Group documents or other text items by semantic similarity without requiring predefined labels.
  • Classification: Use embeddings as features for a downstream classifier or similarity-based categorization system.
  • Multilingual retrieval: Embed text across more than 50 major languages for applications that need cross-language or multilingual search.

For a basic document search workflow, an application can split documents into passages, submit those passages for embedding, and store the returned vectors with the original text and metadata. When a user submits a query, the application embeds the query using the same model and searches for nearby stored vectors. The retrieved passages can then be displayed directly or supplied to a generative model.

Strengths and practical trade-offs

The most important strength of text-embedding-v3 is not text generation or tool use; it is its focused role as a multilingual embedding model. The documented language coverage and 8,192-token input capacity make it suitable for many ordinary search and retrieval pipelines. Its configurable dimensions also give developers a choice between the default representation and smaller vectors that can reduce storage and indexing overhead.

Compatibility is another central reason to use it. Alibaba Cloud describes text-embedding-v3 as useful for maintaining existing v3 vector indexes. If an application already has a production index built with this model, continuing to use the same model can avoid the operational work of re-embedding documents, rebuilding indexes, validating retrieval quality, and coordinating a migration across services.

The trade-off is that the provider recommends text-embedding-v4 for new text-search and RAG deployments. That recommendation does not make text-embedding-v3 unusable, but it does mean that teams starting from scratch should compare the newer option before committing to a v3-based index. The supplied research does not provide a numerical benchmark showing how much one model outperforms the other, so a quality difference should not be assumed without testing the application's own data.

Editorially, this model receives a speed score of 8 out of 10 and a cost score of 8 out of 10 for embedding workloads. These are comparative estimates for this review, not Alibaba Cloud ratings or published benchmark results. They reflect the model's focused embedding role and low listed token price, but actual throughput and total cost depend on request size, rate limits, index storage, database infrastructure, and application design.

Pricing and rate limits

For the Singapore international deployment, the listed price is $0.07 per 1 million input tokens. Output is listed as free. Alibaba Cloud also documents a 500,000-token free quota in the supplied pricing research. Pricing can depend on deployment or region, so users should verify the applicable Model Studio pricing page before production use.

The international rate limits supplied for the model are 6,000 requests per minute and 24,000,000 tokens per minute. These limits are shared operational constraints rather than a guarantee that every request will complete at the same speed. Applications handling large ingestion jobs should batch work within the documented API behavior, monitor throttling, and design retry handling for rate-limit responses.

Because embedding pricing is based on input tokens, document chunking has a direct cost impact. Very small chunks can increase the number of requests and create more vectors to store, while very large chunks can make retrieval less precise and approach the 8,192-token input limit. A practical design should balance retrieval quality, token usage, request volume, and vector-database costs.

Limitations and unsupported capabilities

text-embedding-v3 accepts text and returns embeddings. It is not a multimodal embedding system: the supplied specifications do not support image, audio, or video input. It also is not intended for text generation, so it should not be selected as the component responsible for writing chatbot responses, producing summaries, or following multi-step instructions.

The documented capability flags mark tool use, function calling, web search, structured outputs, context caching, batch inference, and fine-tuning as unsupported. There is therefore no supported function-calling or web-search workflow to attach directly to this model. An application can combine its vectors with other services, databases, or language models, but those capabilities belong to the surrounding system rather than to text-embedding-v3 itself.

Reasoning and coding scores are not provided in the research, and they are not meaningful primary measures for an embedding model. The model should not be evaluated as a reasoning or code-generation model. Similarly, the research gives no separate knowledge-cutoff date; embedding models transform supplied input rather than answering from a published conversational knowledge cutoff.

When to choose text-embedding-v3

Choose text-embedding-v3 when an application already uses v3 vectors and preserving index compatibility is more valuable than changing models. It is also a reasonable fit when the application needs multilingual text embeddings, semantic search, RAG retrieval, recommendations, clustering, or classification and the documented 8,192-token limit and supported dimensions meet its requirements.

It may be particularly practical for an established pipeline where re-embedding a large corpus would create migration risk or operational expense. Keeping the same model for both stored documents and incoming queries also avoids an immediate need to rebuild the index.

Consider text-embedding-v4 instead when building a new text-search or RAG system, because Alibaba Cloud currently recommends it for those new deployments. The supplied research does not include comparative pricing or benchmark measurements, so the decision should be validated with representative documents, queries, languages, latency requirements, and cost estimates. A newer model may be preferable for a new index, while text-embedding-v3 can remain the lower-disruption choice for an existing one.

Bottom line

text-embedding-v3 is a focused, multilingual text-to-vector model rather than a general-purpose conversational AI system. Its documented 8,192-token input capacity, support for more than 50 major languages, configurable vector dimensions, and low international input price make it suitable for common retrieval and similarity workloads. Its clearest role in Alibaba Cloud's current catalog is compatibility: it remains useful for existing v3 embedding indexes, while new projects should evaluate the provider-recommended text-embedding-v4 before creating a long-lived index.


Answers to Frequently Asked Questions

What is text-embedding-v3 used for?
text-embedding-v3 converts text into numerical vectors for semantic search, retrieval-augmented generation (RAG), recommendations, clustering, classification, and multilingual retrieval. It is designed to represent text meaning rather than generate conversational answers.
Which languages and input sizes does text-embedding-v3 support?
text-embedding-v3 supports more than 50 major languages and accepts up to 8,192 tokens per synchronous request. Longer documents should be split into smaller passages before they are embedded.
What vector dimensions are available with text-embedding-v3?
The default output size is 1,024 dimensions, with 768- and 512-dimensional options also available. Smaller vectors can reduce storage and similarity-search costs, but changing dimensions for an existing index generally requires an indexing or re-embedding plan.
How much does text-embedding-v3 cost and what are its rate limits?
For the Singapore international deployment, the listed price is $0.07 per 1 million input tokens, with output listed as free and a documented 500,000-token free quota in the supplied pricing research. The international limits are 6,000 requests per minute and 24,000,000 tokens per minute. Pricing and limits may vary by deployment or region.
Should I choose text-embedding-v3 or text-embedding-v4?
text-embedding-v3 is a practical choice for maintaining existing v3 vector indexes because it avoids re-embedding documents and rebuilding retrieval infrastructure. For new text-search or RAG deployments, Alibaba Cloud recommends evaluating text-embedding-v4. Teams should compare the models using representative data, languages, quality requirements, latency targets, and costs.


Sources 5
Provider

About Qwen