Rerank 3.0

rerank-multilingual-v3.0

by Cohere · Current older-generation model; superseded by newer Rerank model generations

Cohere Rerank Multilingual v3.0 is a specialized second-stage retrieval model that scores and reorders candidate documents against a query across more than 100 languages. It supports semi-structured data, uses a 4,096-token context window, and returns ranking information rather than generated answers.

Reasoning Coding
Cohere Rerank Multilingual v3.0 improves an existing search pipeline by examining the relationship between a query and a shortlist of candidate documents, then returning those candidates in relevance order. Its multilingual coverage, 4,096-token context limit, support for semi-structured data, and normalized relevance scores make it useful for enterprise search, retrieval-augmented generation, question answering, and cross-language information retrieval.
Inputs

What it can understand

Text
Model profile

Performance characteristics

5/10 Reasoning
4/10 Coding
7/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Rerank 3.0
Model type Other
Context window 4K tokens
Maximum output tokens
Release date 2024-04-09
Status Current older-generation model; superseded by newer Rerank model generations
Knowledge cutoff notes

A model knowledge cutoff is not applicable or publicly specified for this specialized reranking endpoint. It scores submitted queries and documents rather than providing standalone generative answers.

Model notes

The canonical API model identifier is rerank-multilingual-v3.0. It reranks queries against documents and semi-structured data rather than generating natural-language answers. Cohere documents support for more than 100 languages and a 4,096-token context length. Query and document tokens contribute to the context calculation, and long documents may be automatically chunked. Rerank usage is billed by search units rather than input and output generation tokens; an exact current public self-serve price for this model was not verified. The current Cohere catalog also lists newer Rerank 3.5 and Rerank 4 models.

Model guide

Cohere Rerank Multilingual v3.0 for Cross-Language Search and RAG

Cohere Rerank Multilingual v3.0 is a specialized semantic reranking model that reorders search results, documents, and semi-structured records by relevance to a query across more than 100 languages. It is designed for second-stage retrieval rather than conversation or text generation.

What Cohere Rerank Multilingual v3.0 does

Cohere Rerank Multilingual v3.0 is a semantic reranking model from Cohere. It is not a chatbot, embedding model, or general-purpose text generator. Instead, it performs a focused job in a search pipeline: it compares a user query with a set of candidate documents and estimates how relevant each candidate is to that query.

A typical system first retrieves a manageable shortlist using keyword search, vector search, or a combination of both. Rerank Multilingual v3.0 then examines the query-document pairs more directly and rearranges the shortlist. This second-stage process can improve the position of the most useful results when the initial search system returned several plausible but imperfect matches.

The model was released on April 9, 2024. In Cohere's current catalog it is an older-generation Rerank model: the provider now also lists newer Rerank 3.5 and Rerank 4 models. That positioning does not make v3.0 unsuitable for existing workloads, but teams starting a new deployment should compare it with those newer options for quality, language coverage, latency, and availability.

Core capabilities and supported data

The model supports multilingual semantic reranking across more than 100 languages, according to Cohere's documentation and model information. This makes it suitable for collections where documents are written in multiple languages or where the query and result language may differ. For example, a user could search a multilingual knowledge base using a query in one language and use the reranker to help identify relevant material written in another.

  • Query-document relevance: compares each candidate with the submitted query rather than relying only on keyword overlap.
  • Multilingual retrieval: supports more than 100 languages according to the supplied Cohere documentation.
  • Semi-structured data: can rerank records represented as JSON or similar text formats, which is useful for product catalogs, support records, and business databases.
  • Normalized scores: returns relevance scores on a 0-to-1 range, allowing applications to sort results or apply an application-specific threshold.
  • Automatic chunking: Cohere documents automatic handling of long query-document combinations when they exceed the applicable context limit.

Reranking does not create a final answer for the user. Its output is ranking information associated with the submitted documents. A separate retrieval application or generative model must use those results to display documents, construct citations, or write an answer.

Context length and input limits

Rerank Multilingual v3.0 has a 4,096-token context length for the combined query and document processing window. The query and each document both contribute to the context calculation. This is important when sending long records: a document that appears short in characters may still consume many tokens, particularly when it contains dense text or structured content.

Cohere's reranking documentation states that long documents may be automatically chunked. Automatic chunking reduces the need for an application to split every document manually, but it does not remove the need for sensible retrieval design. Sending an entire corpus, very large records, or an unnecessarily large candidate set can increase processing work and make the ranking pipeline less efficient.

The model does not have a conventional maximum generated-output limit because it does not generate prose. Its response consists of relevance and ranking information for the submitted candidates. The supplied model record does not specify a separate maximum-output-token value.

Where it fits in a retrieval pipeline

Rerank Multilingual v3.0 is best understood as a second-stage retrieval component. A practical pipeline may work as follows:

  1. A user submits a natural-language query.
  2. A keyword, vector, or hybrid search system retrieves a candidate set.
  3. Rerank Multilingual v3.0 compares the query with those candidates.
  4. The application sorts the candidates by the returned relevance scores.
  5. The highest-ranked passages are shown to the user or supplied to a separate answer-generation model.

This division of labor is useful because initial retrieval systems are optimized for quickly finding possible matches, while a reranker can spend more computation evaluating the relationship between the query and each candidate. The reranker therefore does not replace an index or database search engine. It refines the results produced by one.

For retrieval-augmented generation, or RAG, the model can improve the evidence selection stage before a generative model writes an answer. Better-ranked passages may reduce the chance that the answer model receives irrelevant or weakly related context, although the reranker itself does not verify facts or produce citations.

Strengths and trade-offs

The clearest strength of this model is its specialization. It is focused on relevance ranking rather than spending resources on capabilities that are unnecessary for a search pipeline. Its multilingual coverage is particularly relevant to international support portals, enterprise document repositories, and cross-language search applications.

Its ability to process semi-structured data is another practical advantage. Instead of limiting reranking to paragraphs of prose, an application can represent fields from a business record in a text or JSON-like form and ask the model to judge the record against a query. The quality of that approach depends on how clearly the fields are represented and which fields are included.

The main trade-off is that it is not a complete AI assistant. It cannot answer a question, summarize a document, write code, create an image, or generate audio and video. It also does not provide native multimodal input or output in the supplied model specification: its supported input is text, and its output is ranking data. Applications needing a final natural-language response must add another component.

Reranking also introduces an extra stage and an additional cost after initial retrieval. In return, it can improve the quality of a relatively small candidate set. The appropriate balance depends on the application's latency and budget requirements. A system that needs the lowest possible latency may use only its initial search method, while a system where relevance quality matters more may rerank the top candidates.

Pricing and access

Cohere exposes Rerank Multilingual v3.0 through its Rerank API and selected partner deployment environments. The model's usage is billed by search units rather than by generated output tokens or the usual input-and-output token arrangement used for generative models.

An exact current public self-serve price for this specific model was not verified in the supplied research. Consequently, there is no reliable numeric price to report here. Teams should check Cohere's current pricing and account documentation before estimating production costs, particularly because pricing may depend on the deployment environment and the number of searches or documents processed.

The search-unit billing model means that cost planning should account for how many candidate documents are sent to the reranker for each query. Reducing the candidate set before reranking can affect both latency and usage, but an overly small set may exclude the document that should have ranked first.

Modalities, reasoning, coding, and tools

Rerank Multilingual v3.0 accepts textual queries and textual representations of documents or semi-structured records. It returns scores and ranking information rather than direct non-text media or generated prose. It therefore has no image, audio, or video input or output capability in the supplied specification.

Terms such as reasoning and coding are less applicable to this endpoint than they are to a conversational language model. The model performs semantic relevance assessment, but it is not intended to show a chain of reasoning or solve programming tasks. It should not be selected for code generation, code execution, tool calling, or function orchestration. The supplied model record marks tool use, streaming, JSON mode, caching, batch API support, and fine-tuning as unavailable or unsupported for this model.

Those limitations describe the model endpoint, not every feature that might exist in a larger Cohere deployment or application. A surrounding system may combine the reranker with search infrastructure, business logic, or another model, but those added capabilities should not be attributed to Rerank Multilingual v3.0 itself.

Best use cases

  • Multilingual enterprise search: reorder internal policies, procedures, and knowledge-base articles for employees in different regions.
  • Customer-support retrieval: identify the most relevant support articles or prior cases before presenting them to an agent or answer-generation system.
  • RAG evidence selection: select stronger passages for a separate generative model to use as context.
  • Cross-language information retrieval: improve search when the query and candidate documents are not all written in the same language.
  • Document discovery: rank contracts, reports, manuals, or records against a natural-language request.
  • Structured business search: compare a query with product, customer, ticket, or inventory records represented as text or JSON-like data.

When to choose this model

Choose Rerank Multilingual v3.0 when an existing search system already produces candidates but their ordering is not reliable enough, especially when the corpus spans many languages or includes semi-structured records. It is a good fit when the application needs relevance scores and ranked documents rather than a conversational response.

It may be preferable to a general-purpose generative model when the task is only ranking. Using a specialized reranker avoids asking a chat model to imitate a scoring system and keeps the output focused on retrieval decisions. It can also be more appropriate than relying on vector similarity alone when exact semantic matching between a query and each candidate matters.

Consider a newer Cohere Rerank generation, including Rerank 3.5 or Rerank 4, when beginning a new implementation and current model availability, quality, language support, or operational characteristics are more important than compatibility with an existing v3.0 integration. The supplied research does not establish a detailed benchmark comparison, so those newer models should be evaluated rather than assumed to be better for every workload.

Choose another type of model when the application must generate an answer, summarize or transform text, write code, interpret images, or call tools. Rerank Multilingual v3.0 can help select the information needed for those tasks, but it is not the component that performs them.

Bottom line

Cohere Rerank Multilingual v3.0 is a focused multilingual relevance model for improving the order of search results. Its more-than-100-language coverage, 4,096-token context length, support for semi-structured data, and 0-to-1 relevance scores make it useful in enterprise retrieval and RAG pipelines. Its boundaries are equally important: it does not generate answers, provide multimodal output, or replace the search system that supplies its candidates. It is most valuable as a disciplined second stage between retrieval and whatever application ultimately presents or generates the result.


Answers to Frequently Asked Questions

How does Cohere Rerank Multilingual v3.0 fit into a RAG pipeline?
In a RAG pipeline, an initial keyword, vector, or hybrid search retrieves candidate passages. Rerank Multilingual v3.0 then scores and reorders those candidates so that a separate generative model receives more relevant context. The reranker itself does not verify facts or produce citations.
What are the context length and input limits of Rerank Multilingual v3.0?
Rerank Multilingual v3.0 has a 4,096-token context length for the combined query and document processing window. Cohere documents automatic chunking for long query-document combinations, but applications should still use sensible document sizes and candidate sets.
Does Cohere Rerank Multilingual v3.0 generate answers?
No. The model returns relevance scores and ranking information for submitted candidates rather than generating prose or final answers. A separate application or generative model must display the results, create citations, or write a response.
What is Cohere Rerank Multilingual v3.0 used for?
Cohere Rerank Multilingual v3.0 is used as a second-stage retrieval component that compares a query with candidate documents and reorders them by semantic relevance. It is suitable for multilingual search, enterprise retrieval, cross-language information retrieval, and RAG evidence selection.
How many languages does Cohere Rerank Multilingual v3.0 support?
According to Cohere’s documentation and model information, Rerank Multilingual v3.0 supports more than 100 languages. This makes it useful for multilingual collections and searches where the query and documents may be written in different languages.


Sources 7
Provider

About Cohere