Qwen3-Embedding

text-embedding-v4

by Qwen · Current and accessible through Alibaba Cloud Model Studio

Alibaba Cloud text-embedding-v4 is a Qwen3-Embedding model for multilingual semantic search, RAG, document retrieval, clustering, classification, recommendation, and code retrieval. It supports more than 100 languages, inputs up to 8,192 tokens, batches of up to 10 texts, and vector dimensions from 64 to 2,048. Dense, sparse, and combined outputs are available, with pricing based on input tokens.

Embeddings Reasoning Coding
Alibaba Cloud text-embedding-v4 is designed for applications that need to compare the meaning of text rather than generate text. It supports more than 100 major languages and multiple programming languages, accepts up to 8,192 tokens per input, and offers configurable vector sizes from 64 to 2,048 dimensions. Its low token price and embedding-specific batch processing make it a practical option for large search and retrieval workloads.
Outputs

What text-embedding-v4 can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

1/10 Reasoning
2/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen3-Embedding
Model type Other
Context window 8K tokens
Status Current and accessible through Alibaba Cloud Model Studio
Knowledge cutoff notes

Alibaba Cloud's official documentation does not publish a knowledge cutoff for this embedding model. As an embedding endpoint, it produces vector representations rather than current factual answers.

Model notes

Developed by Tongyi Lab and described by Alibaba Cloud as part of the Qwen3-Embedding series. It supports more than 100 major languages and multiple programming languages. Supported dimensions are 64, 128, 256, 512, 768, 1024, 1536, and 2048, with 1024 as the default. Each request supports up to 10 texts, with each text limited to 8192 tokens. The API can return dense, sparse, or dense-and-sparse vectors; sparse output is supported through the DashScope API. Alibaba Cloud documentation lists batch processing for large-scale embedding workloads, while the model capability table marks general batch inference as unsupported, so batch support should be understood as embedding-specific batch API/file processing rather than general generative batch inference. No official knowledge cutoff or exact release date was located.

Cost

Model pricing

Input $0.07 per 1M input tokens in Singapore/International and Hong Kong; $0.072 per 1M input tokens in China (Beijing). Beijing batch-file processing is listed at $0.036 per 1M tokens.
Output Free; pricing is based on input tokens
Model guide

Alibaba Cloud text-embedding-v4 for Multilingual Search and RAG

Alibaba Cloud text-embedding-v4 is a multilingual Qwen3-Embedding model that converts text into dense, sparse, or combined vector representations for semantic search, retrieval-augmented generation, clustering, classification, recommendation, and code retrieval.

What is Alibaba Cloud text-embedding-v4?

Alibaba Cloud text-embedding-v4 is a text embedding model in the Qwen3-Embedding family, developed by Tongyi Lab and available through Alibaba Cloud Model Studio. Instead of returning an answer in natural language, it converts text into numerical vectors. These vectors represent patterns of meaning, allowing software to identify semantically similar documents, queries, passages, code samples, or categories.

For example, a search system can compare a user query such as “how do I reset my password?” with documents that use different wording, such as “account credential recovery instructions.” A vector search engine can recognize that the two texts are related even when they do not share many exact keywords.

The model is therefore intended for retrieval and analysis workflows rather than conversation. It can provide the vector representation used by a semantic search system, a retrieval-augmented generation (RAG) pipeline, a recommendation engine, or a clustering and classification process.

Where it fits in Alibaba Cloud's model catalog

text-embedding-v4 is listed as a current model accessible through Alibaba Cloud Model Studio. Its Qwen3-Embedding lineage places it among Alibaba's embedding-focused models rather than its text-generation models. The endpoint is specialized: its output is an embedding representation, not a conversational response or generated document.

This distinction matters when selecting a model. An embedding model is usually one component inside a larger application. A typical RAG system might use text-embedding-v4 to index documents and find relevant passages, then pass those passages to a separate language model that generates the final answer. text-embedding-v4 itself is not the component that writes that answer.

Core capabilities and input limits

SpecificationVerified detail
Model familyQwen3-Embedding
ProviderAlibaba Cloud
Input typeText
Output typeDense, sparse, or dense-and-sparse vectors
LanguagesMore than 100 major languages, plus multiple programming languages
Maximum input length8,192 tokens per text
Supported dimensions64, 128, 256, 512, 768, 1,024, 1,536, or 2,048
Default dimension1,024
Texts per requestUp to 10 texts

The 8,192-token limit applies to each text input. Applications processing long documents generally need to split them into smaller passages before creating embeddings. Chunking is an application-level decision: very small chunks may lose context, while very large chunks can make retrieval less precise and may approach the model's input limit.

The selectable vector dimension provides a practical storage and performance trade-off. A 2,048-dimensional vector retains a larger representation but requires more storage and can increase the workload for a vector database. Smaller options such as 256 or 512 dimensions can reduce storage and similarity-search costs, although the research supplied here does not establish comparative quality results for each dimension. The 1,024-dimensional setting is the documented default.

Dense, sparse, and combined embeddings

text-embedding-v4 can return dense vectors, sparse vectors, or both together. A dense vector stores a fixed-length numerical representation in which many positions contain values. Sparse representations emphasize selected features and can complement semantic matching with more selective signals.

According to Alibaba Cloud's documentation, sparse output is supported through the DashScope API. The choice between output types should be based on the search architecture and vector database being used. Dense output is a straightforward fit for conventional vector similarity search. Combined dense-and-sparse output may be useful when an application wants to use both semantic similarity and more selective lexical or feature-level matching, but the exact retrieval design remains an implementation choice.

What is text-embedding-v4 best used for?

  • Semantic search: Find relevant content by meaning, including queries and documents that use different wording.
  • RAG retrieval: Convert a document collection into vectors and retrieve relevant passages before a separate generative model produces an answer.
  • Document retrieval: Search internal knowledge bases, manuals, policies, support content, or other text collections.
  • Clustering: Group related documents, messages, tickets, or other text records based on their vector representations.
  • Classification: Use embeddings as features for assigning text to categories.
  • Recommendation: Compare the meaning of user interests, documents, products, or content descriptions.
  • Code retrieval: Search programming content using natural-language descriptions or related code concepts.

Its multilingual support is particularly relevant to applications that index content or receive queries in multiple languages. The supplied research verifies support for more than 100 major languages and multiple programming languages, but it does not provide a language-by-language quality ranking or benchmark.

Pricing and batch processing

Alibaba Cloud's listed pricing is based on input tokens. In Singapore, International, and Hong Kong, the researched price is $0.07 per 1 million input tokens. In China (Beijing), the listed price is $0.072 per 1 million input tokens. Beijing batch-file processing is listed separately at $0.036 per 1 million tokens. Output is free because the service returns embeddings rather than priced generated text.

Region or processing modePrice
Singapore, International, and Hong Kong$0.07 per 1 million input tokens
China, Beijing$0.072 per 1 million input tokens
Beijing batch-file processing$0.036 per 1 million tokens
OutputFree

These are token-based usage prices, not a monthly subscription fee. Actual deployment costs can also include storage, vector-database queries, networking, and any separate model used to generate answers in a RAG system. Alibaba Cloud documentation describes batch processing for large-scale embedding workloads. The supplied research also notes an important qualification: the general capability table marks broad batch inference as unsupported, so batch support should be understood as embedding-specific batch API or file processing rather than a general-purpose generative batch-inference feature.

Speed, cost, and capability trade-offs

Embedding endpoints are generally used for high-volume transformations: a large document collection may need to be indexed once, and new queries may need to be embedded continuously. text-embedding-v4's low listed input-token price and support for batch processing are its clearest operational advantages for this type of workload.

In the supplied editorial evaluation, the model receives a speed score of 8 out of 10 and a cost score of 9 out of 10. These are editorial assessments, not Alibaba Cloud-published benchmark results. They indicate that the model is viewed as a fast, inexpensive option for embedding workloads, but they should not be interpreted as a guaranteed latency or quality measurement.

The main capability trade-off is specialization. text-embedding-v4 is not a text-generation model, so it cannot independently produce a conversational answer, summarize a retrieved document, or follow a multi-step tool-using workflow. A system that needs those functions must combine the embedding endpoint with other components.

Reasoning, coding, tools, and modalities

text-embedding-v4 does not provide a conversational reasoning process or generated reasoning trace. Its role is to encode input text into vectors. The research gives it an editorial reasoning score of 1 out of 10, reflecting that it is not intended for reasoning-based response generation rather than measuring a provider-published reasoning benchmark.

It can be used for code retrieval because the model supports multiple programming languages, but it is not a code-generation or software-engineering agent. The supplied editorial coding score is 2 out of 10, again indicating limited direct coding capability rather than a published coding benchmark.

The verified input modality is text, and the verified output is an embedding representation. It does not support image, audio, or video input, and it does not directly output images, audio, video, or natural-language text. Tool or function calling is not supported as a model capability, and streaming is not relevant to its vector output in the supplied capability information.

Important limitations

  • Not a generator: It cannot answer users directly or create natural-language responses.
  • Text-only: It is not suitable for image, audio, or video understanding or multimodal retrieval.
  • Input-size ceiling: Each text is limited to 8,192 tokens, so long documents require preprocessing and chunking.
  • Request-size ceiling: A synchronous request supports up to 10 texts.
  • Not a reranker: The research identifies reranking as a task for which it is not intended. A separate reranking option may be needed after initial retrieval.
  • No native tool workflow: It does not perform function calls or autonomous actions.
  • No published knowledge cutoff: Alibaba Cloud does not publish a knowledge cutoff for this embedding model. Since it creates representations rather than current factual answers, a knowledge cutoff is less directly applicable than it is for a generative model.

When to choose text-embedding-v4

Choose text-embedding-v4 when the central requirement is affordable, multilingual text representation for search, retrieval, classification, clustering, recommendation, or code-search systems. It is especially suitable when the application needs configurable vector dimensions, support for dense and sparse retrieval designs, or batch processing for a sizeable document collection.

It may be a good fit for a multilingual knowledge base, an internal enterprise search engine, a support-ticket retrieval system, or a RAG pipeline in which another model handles answer generation. Its low input-token price can also make it attractive when many documents or queries must be embedded.

Consider another option when the application needs direct text generation, multimodal input, native reranking, tool use, or a conversational agent in one model. A separate generative model is more appropriate for producing answers, while a dedicated reranker may be preferable when the retrieval pipeline needs a second-stage relevance model. The supplied research does not identify a specific alternative model or provide quality benchmarks against named competitors, so those choices should be tested against the application's languages, documents, latency requirements, and retrieval metrics.

Bottom line

Alibaba Cloud text-embedding-v4 is a focused embedding service rather than an all-purpose AI assistant. Its main strengths are multilingual coverage, support for programming languages, configurable dimensions, multiple vector output formats, an 8,192-token input limit, and low token-based pricing. Its limitations are equally clear: it produces vectors only, does not handle non-text modalities, does not generate answers or use tools, and requires additional retrieval and generation components for a complete RAG or conversational application.


Answers to Frequently Asked Questions

What is Alibaba Cloud text-embedding-v4 used for?
Alibaba Cloud text-embedding-v4 converts text into numerical vectors for semantic search, RAG retrieval, document retrieval, clustering, classification, recommendation, and code retrieval. It is an embedding model, not a conversational text-generation model.
Does text-embedding-v4 support multilingual search?
Yes. text-embedding-v4 supports more than 100 major languages as well as multiple programming languages, making it suitable for multilingual knowledge bases, enterprise search, and cross-language retrieval. The available information does not provide language-by-language quality rankings or benchmark results.
What are the input and vector dimension limits of text-embedding-v4?
Each text input can contain up to 8,192 tokens, and a synchronous request can include up to 10 texts. Supported vector dimensions are 64, 128, 256, 512, 768, 1,024, 1,536, and 2,048, with 1,024 as the default.
How much does Alibaba Cloud text-embedding-v4 cost?
The listed price is $0.07 per 1 million input tokens in Singapore, International, and Hong Kong, and $0.072 per 1 million input tokens in China (Beijing). Beijing batch-file processing is listed at $0.036 per 1 million tokens, and embedding output is free. Additional costs may apply for storage, vector databases, networking, and separate generation models.
Can text-embedding-v4 generate answers or perform reranking?
No. text-embedding-v4 only produces dense, sparse, or combined embeddings. It cannot generate natural-language answers, perform native tool calls, process images, audio, or video, or act as a dedicated reranker. RAG applications typically combine it with a vector database, an optional reranker, and a separate generative model.


Sources 5
Provider

About Qwen