What Cohere Rerank 4 Pro does
Cohere Rerank 4 Pro is a second-stage retrieval model. A search system first finds a set of potentially relevant documents using keyword search, vector search, or a combination of both. Rerank 4 Pro then examines the user's query alongside those candidates and places the most relevant results first.
This distinction matters because the model is not intended to search an entire knowledge base by itself. It does not replace a search index or embedding model, and it does not write a final answer. Instead, it improves the material that a search interface or generative model receives after the initial retrieval step.
For example, an enterprise support system could retrieve 100 documents about a product problem, send the query and those documents to Rerank 4 Pro, and pass the highest-ranked results to an answer-generation model. The reranker helps reduce irrelevant context before the answer-generation stage.
Where it fits in Cohere's lineup
Cohere provides Rerank 4 Pro as part of its Rerank 4.0 model family. The canonical API identifier is rerank-v4.0-pro. Cohere positions the Pro variant for the highest quality and more complex relevance-ranking workloads, while Rerank 4 Fast is intended for applications that prioritize lower latency and higher throughput.
That makes Rerank 4 Pro a specialist component rather than a general-purpose language model. It is focused on judging the relationship between supplied queries and supplied content. Its output consists of relevance scores and ranking information, not generated paragraphs, chat responses, embeddings, images, audio, or video.
Key capabilities and supported content
The model supports English and non-English documents in more than 100 languages, according to Cohere's model information. This makes it suitable for multilingual enterprise search, where users may search across documents written in different languages or where query and document languages do not always match.
Rerank 4 Pro can process ordinary text documents as well as semi-structured data represented in supported formats such as JSON or YAML strings. This is useful when the content being ranked consists of records, metadata-rich entries, product catalogs, structured support cases, or other data that does not read like a conventional prose document.
Its documented context window is 32,768 tokens. A token is a small unit of text used by the model; the context limit describes how much query and document material can be processed in one comparison. Cohere's documentation notes that documents are processed in chunks of up to 32,764 tokens after reserved tokens, and that queries can account for up to half of the context window.
Context and request limits
The query can contain up to 16,384 tokens before truncation. Documents are handled within the model's 32,768-token context, and long documents may be divided automatically into chunks. The endpoint supports up to 10,000 documents per request under the documented document and chunking constraints.
These limits should not be interpreted as a guarantee that every large request will have identical latency or cost. A request containing many long documents can require substantially more processing than a small set of short passages. Applications should also check how chunking affects the ranking of very long documents, particularly when important evidence is spread across multiple sections.
Pricing and availability
Cohere released Rerank 4.0 on December 11, 2025. Public API usage for Rerank 4 Pro is priced at $2.50 per 1,000 search units. A search unit represents one query against up to 100 documents. This is a ranking price, not a charge for generated output tokens, because the endpoint does not generate an answer.
The effective cost depends on how an application groups documents into requests and how many queries it sends. For example, sending one query against up to 100 documents is treated differently from repeatedly reranking smaller groups for the same user task. Teams should therefore estimate usage from their actual retrieval pattern rather than comparing the headline rate directly with text-generation prices.
Cohere also offers dedicated Model Vault capacity for Rerank 4 Pro. Model Vault pricing is based on provisioned instance size and is separate from public API search-unit pricing. The supplied research does not establish a single universal Model Vault price because the cost depends on the selected capacity and deployment arrangement.
Reasoning, coding, and output behavior
Rerank 4 Pro performs relevance reasoning in the narrow sense required to compare a query with candidate content. It can judge which supplied documents are more relevant, but it is not a general reasoning model for solving multi-step problems or explaining an answer.
Its coding capability is similarly limited. The model can rank code-related documents, technical records, or snippets when they are supplied as candidate content, but it is not intended to generate, debug, or execute software. The research also indicates no tool or function-calling support, no streaming output, and no separate structured-output mode.
The output is text-based ranking information, typically including relevance scores and result ordering. There is no maximum generated-token setting because the endpoint is not a text-completion service. It also has no image, audio, or video input or output capability.
Main strengths and trade-offs
- Multilingual retrieval: support for more than 100 languages is useful for international knowledge bases and cross-language enterprise search.
- Quality-focused positioning: Pro is intended for difficult relevance decisions where ranking quality matters more than maximum throughput.
- Long comparisons: the 32,768-token context window supports longer query-document comparisons than Cohere's older 4,096-token Rerank models.
- Structured content: JSON and YAML strings can be ranked alongside ordinary text, helping applications search records and metadata-rich content.
- Large candidate sets: requests can include up to 10,000 documents under the documented chunking constraints.
The principal trade-off is speed and cost. Rerank 4 Pro is less appropriate than Rerank 4 Fast when an application needs the lowest possible latency or the highest throughput. The Pro option is most defensible when better ordering of retrieved content has enough value to justify additional processing.
Best use cases
- Retrieval-augmented generation: rerank passages before sending them to a generative model, improving the chance that the model sees the most relevant evidence.
- Multilingual enterprise search: rank internal policies, support material, or knowledge-base content across many languages.
- Customer-support retrieval: prioritize articles and historical cases that best match a user's problem.
- Structured record search: rank product records, tickets, catalog entries, or JSON and YAML documents using natural-language queries.
- Hybrid search: improve the final ordering after keyword and vector retrieval have produced a candidate set.
- High-precision retrieval: use a quality-oriented second stage before passing limited context to an answer-generation model.
When to choose Rerank 4 Pro
Choose Rerank 4 Pro when the main problem is not generating text but deciding which retrieved items deserve priority. It is a strong fit for multilingual or enterprise search systems where irrelevant context can reduce answer quality, and for applications that need to compare long or semi-structured documents.
Choose Rerank 4 Fast instead when latency, request volume, or throughput is more important than the highest available ranking quality within Cohere's Rerank 4.0 family. Choose an embedding model when you need to create vector representations for initial retrieval. Choose a generative model when the system must explain findings, answer questions, summarize documents, or produce code.
Rerank 4 Pro is therefore best understood as a precision layer in a larger retrieval pipeline. It can make an existing search or RAG system more selective, but it cannot operate as a complete search application or conversational assistant on its own.
Limitations to consider
The model requires both a query and candidate documents, so it has little value without an upstream retrieval process or a defined set of items to compare. It does not independently crawl the web, build an index, create embeddings, or provide final answer generation.
Its Pro quality tier may also be excessive for simple ranking tasks or applications where response time is critical. Long documents can be chunked, and ranking many large candidates can increase processing requirements. Finally, Cohere's documentation states that fine-tuning for Rerank models was retired in a September 16, 2025 deprecation announcement, so organizations should not plan on customizing this model through the former fine-tuning workflow.
Bottom line
Cohere Rerank 4 Pro is a specialized, multilingual relevance model for improving the ordering of retrieved content. Its 32,768-token context, support for more than 100 languages, semi-structured data handling, and quality-focused positioning make it suitable for demanding enterprise search and RAG pipelines. Its limitations are equally important: it does not generate answers, provide embeddings, use tools, or replace the rest of a retrieval system. The practical choice is straightforward: use it when ranking quality is the bottleneck, and consider the faster Rerank 4 Fast option when throughput and latency matter more.

