Qwen3 Reranker

qwen3-rerank

by Qwen · Current and available

Alibaba Cloud's managed multilingual text reranking model for improving search and RAG results. Qwen3-Rerank supports more than 100 languages, up to 500 documents per request, 4,000 input tokens per item, and 120,000 total input tokens. The international Singapore deployment costs $0.10 per one million input tokens, with output free.

Reasoning Coding
Qwen3-Rerank is a specialized model for improving search quality rather than generating conversational answers. It takes a query and candidate text passages, evaluates their relevance, and returns a better ordering for search, enterprise knowledge bases, and retrieval-augmented generation pipelines.
Inputs

What it can understand

Text
Model profile

Performance characteristics

2/10 Reasoning
2/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen3 Reranker
Model type Other
Context window 4K tokens
Maximum output tokens
Release date 2025-06-05
Status Current and available
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for the managed reranking API. Reranking operates on the query and candidate documents supplied at request time.

Model notes

The canonical managed API model ID is qwen3-rerank. Alibaba Cloud documents more than 100 supported languages, a maximum of 500 documents per request, 4,000 input tokens per query or document item, and 120,000 total input tokens per request. The API returns relevance-ranking results rather than normal generated prose. Alibaba Cloud's documentation does not state an exact parameter count for the managed API model. Separately, the Qwen team released open-weight Qwen3-Reranker checkpoints in 0.6B, 4B, and 8B sizes in June 2025; those checkpoints should not automatically be assumed to be identical to the managed API deployment.

Cost

Model pricing

Input $0.10 per 1 million input tokens for the international Singapore deployment; pricing is region-dependent. Output is free.
Output Free
Model guide

Qwen3-Rerank: Multilingual Search Reranking for Better RAG Results

Qwen3-Rerank is Alibaba Cloud Model Studio's managed text reranking model. It improves search and retrieval-augmented generation by rescoring a shortlist of candidate documents against a query, supporting more than 100 languages, up to 500 documents per request, and 4,000 input tokens per query or document item. The international Singapore deployment is priced at $0.10 per one million input tokens, with no separate output charge.

What is Qwen3-Rerank?

Qwen3-Rerank is Alibaba Cloud Model Studio's managed text reranking model. Its job is to improve the order of results produced by a search or retrieval system. Instead of generating an answer, it compares a user's query with a shortlist of candidate documents and assigns relevance scores so that the most useful passages appear first.

This makes it a second-stage retrieval component. A typical search or retrieval-augmented generation (RAG) workflow first uses a keyword search engine, vector database, or embedding-based retriever to find a broad set of possible matches. Qwen3-Rerank then examines that smaller candidate set in more detail. The best-ranked passages can be passed to a separate language model for answer generation.

The managed API uses the canonical model ID qwen3-rerank. It is a text-only reranker and should not be confused with qwen3-vl-rerank, a separate multimodal model intended for ranking image and video candidates.

How the model fits in a search pipeline

Reranking is useful because the first retrieval stage is usually optimized for speed and recall: it tries to find enough potentially relevant material without examining every document in depth. That initial result set may contain passages that share keywords with the query but do not really answer it.

Qwen3-Rerank provides a more focused relevance check. For example, a company knowledge-base search might initially return documents containing the words “expense policy.” The reranker can help place the passage about international travel reimbursement above a general finance glossary or an outdated policy page when the user's query specifically concerns overseas expenses.

  1. A user submits a query.
  2. An existing search engine or retriever returns candidate passages.
  3. Qwen3-Rerank scores the query against those candidates.
  4. The application keeps or displays the highest-ranked results.
  5. An optional generation model uses the selected passages to produce a final response.

Because the model is designed for ranking rather than final response generation, it does not replace a chat model, an embedding model, or a document database. It complements those components by improving the relevance of the material they retrieve.

Capabilities and supported inputs

Qwen3-Rerank accepts a text query and text documents. Alibaba Cloud documents support for more than 100 major languages, including Chinese, English, Spanish, French, Portuguese, Indonesian, Japanese, Korean, German, and Russian. This multilingual coverage makes it relevant to cross-language search and international knowledge bases, although the supplied documentation does not provide a benchmark comparison for each language.

The model can also accept optional custom instructions. These instructions can help define what “relevant” means for a particular application, such as prioritizing policy documents that apply to a specific region or favoring passages that directly answer a question instead of merely mentioning its subject.

SpecificationDocumented value
Model typeText reranking model
Canonical model IDqwen3-rerank
Supported contentText queries and text documents
Language coverageMore than 100 languages, according to Alibaba Cloud
Maximum query or document item4,000 input tokens
Maximum documents per request500
Maximum total input per request120,000 input tokens
OutputRelevance-ranking results

The 4,000-token limit applies to each query or document item, while the request-level limit is 120,000 input tokens. Applications handling long documents therefore need to split them into passages before reranking. Splitting also makes the resulting ranking more useful because the system can identify the specific section relevant to a query rather than ranking an entire large file as one item.

What does Qwen3-Rerank return?

The API returns ranking information rather than normal generated prose. In practical terms, the application receives relevance results that can be used to reorder the submitted candidates. It is not intended to write a summary, answer a question, create code, or produce a user-facing explanation.

This distinction matters when designing an application. A reranker can identify the strongest supporting passages for a question such as “What is the warranty period for this product?” A separate generation step is still needed if the application should turn those passages into a natural-language answer. If the desired result is a similarity vector for a large-scale first-pass search, an embedding model is a more appropriate component.

Pricing and access

Alibaba Cloud lists Qwen3-Rerank at $0.10 per one million input tokens for the international Singapore deployment. Output is free, and the documented pricing is based on input tokens. Pricing is region-dependent, so users should verify the applicable rate for their deployment before estimating production costs.

The model is accessed through Alibaba Cloud Model Studio's rerank API and requires a Model Studio API key. The supplied research does not establish a separate free tier, guaranteed throughput level, or fixed latency target for this model. Those details should not be assumed from the low input-token price.

At the published Singapore rate, the direct model charge can be attractive for applications that rerank relatively small candidate sets. However, total system cost can also include the initial search service, storage, network usage, application infrastructure, and any separate model used to generate final answers. A cost comparison should therefore evaluate the complete retrieval pipeline rather than the reranker alone.

Main strengths and trade-offs

Qwen3-Rerank's clearest strength is specialization. It is designed for relevance ordering, so it can be inserted between a fast first-stage retriever and a final answer-generation model. Its documented multilingual coverage is another advantage for applications that search content in several languages or serve users across regions.

  • Broad language coverage: Alibaba Cloud claims support for more than 100 languages.
  • Useful request capacity: A request can contain up to 500 documents, subject to the 120,000-token total input limit.
  • Low published input price: The international Singapore rate is $0.10 per million input tokens, with output free.
  • RAG suitability: It can improve the evidence selection stage before a separate model generates an answer.
  • Instruction control: Optional custom instructions can help tailor the ranking policy.

Its limitations are equally important. It does not generate final answers, does not provide image, audio, or video understanding, and is not an embedding generator. It also has no documented tool or function-calling role in the supplied research. The API documentation does not identify an exact parameter count for the managed deployment, so the open-weight model sizes should not be used as a specification for the hosted API.

Managed API and open-weight lineage

The Qwen team released open-weight Qwen3-Reranker models in 0.6B, 4B, and 8B sizes in June 2025 as part of the Qwen3 Embedding and Reranking series. These checkpoints are relevant to users considering local deployment or more direct control over infrastructure.

However, the managed qwen3-rerank API should not automatically be treated as one of those exact checkpoints. Alibaba Cloud's API documentation does not state a parameter-count mapping between the hosted model and the open-weight releases. The hosted service and local checkpoints may therefore differ in deployment behavior, resource requirements, operational controls, or model configuration.

For a managed service, the API is the simpler option when an application needs an available endpoint and does not want to operate model infrastructure. A local open-weight checkpoint may be more appropriate when deployment control, offline operation, or infrastructure customization is more important than managed access. The supplied research does not provide a direct quality, latency, or cost benchmark between these options.

Reasoning, coding, and modality profile

Qwen3-Rerank should not be evaluated like a general-purpose reasoning or coding model. Its task is to calculate relevance between supplied text inputs. The model does not offer conversational reasoning traces, code generation, image generation, audio processing, video processing, or multimodal ranking in the documented qwen3-rerank API.

Its input is text only, and its output is ranking data rather than text content. Tool use, streaming, fine-tuning, caching, batch API support, structured output, and maximum generated-output tokens are not established in the supplied research. In particular, the absence of a documented generated-token limit reflects the model's ranking-oriented output, not a promise that every API feature is unavailable in every Alibaba Cloud interface.

Best use cases

  • RAG document selection: Improve which passages are supplied to a separate answer-generation model.
  • Enterprise knowledge search: Reorder policy, support, product, or internal documentation after an initial retrieval step.
  • Multilingual search: Rank content across the many languages covered by Alibaba Cloud's provider claim.
  • Semantic search: Improve results when keyword or vector retrieval produces a mixed-quality shortlist.
  • Code and text retrieval: Rerank textual documentation, tickets, specifications, or code-related passages when relevance ordering is the main requirement.

It is less suitable when the application needs an all-in-one assistant, direct answer generation, embeddings for indexing, or ranking of images and video. In those cases, a generation model, embedding model, or multimodal reranker may be a better fit.

When to choose Qwen3-Rerank

Choose Qwen3-Rerank when an existing retrieval system already produces candidates but the ordering is not reliable enough. It is especially compelling when multilingual text search, a managed Alibaba Cloud endpoint, and low published input-token pricing are important.

Choose a different type of model when the problem occurs earlier or later in the pipeline. An embedding model is better for creating vectors and performing the initial broad search. A general language model is better for writing the final answer or carrying out a conversation. A multimodal reranker is more appropriate when candidate documents include images or video. A local open-weight reranker may be preferable when infrastructure control or offline deployment is a primary requirement.

Overall, Qwen3-Rerank is best understood as a focused relevance layer. Its value comes from improving the quality of retrieved evidence, not from replacing the rest of a search or RAG system.


Answers to Frequently Asked Questions

What languages and input limits does Qwen3-Rerank support?
Alibaba Cloud states that Qwen3-Rerank supports more than 100 major languages. Each query or document can contain up to 4,000 input tokens, requests can include up to 500 documents, and the total input per request is limited to 120,000 tokens.
What is Qwen3-Rerank used for?
Qwen3-Rerank is a text reranking model that improves the order of search or retrieval results. It scores a user query against a shortlist of candidate documents so the most relevant passages can be shown first or passed to a separate language model in a RAG pipeline.
Does Qwen3-Rerank generate answers or embeddings?
No. Qwen3-Rerank returns relevance-ranking results rather than generated prose, summaries, or conversational answers. It is also not an embedding model, so a separate embedding or search system is needed for first-stage retrieval and a separate language model is needed for answer generation.
How does Qwen3-Rerank fit into a RAG pipeline?
A search engine, vector database, or embedding retriever first finds candidate passages. Qwen3-Rerank then evaluates and reorders those candidates, after which the highest-ranked passages can be supplied to a language model for final answer generation.
How much does the Qwen3-Rerank API cost?
Alibaba Cloud lists the international Singapore deployment at $0.10 per one million input tokens, with output free. Pricing varies by region, and total costs may also include search, storage, network, infrastructure, and any separate answer-generation model.


Sources 7
Provider

About Qwen