Rerank

Rerank 3.5

by Cohere · Active

Cohere Rerank 3.5 is a dedicated second-stage retrieval model that scores and reorders documents against a query. It supports more than 100 languages, YAML-formatted structured records, a 4,096-token evaluation context, and up to 10,000 documents per request subject to service limits. It is designed for enterprise search and RAG rather than conversation or content generation.

Reasoning Coding
Cohere Rerank 3.5 is a second-stage search model for applications that already have a set of candidate documents. Given a query and those candidates, it scores their semantic relevance and returns them in a more useful order. The model supports more than 100 languages, structured records represented as YAML strings, and a 4,096-token context window for each query-document evaluation.
Inputs

What it can understand

Text
Model profile

Performance characteristics

3/10 Reasoning
2/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Rerank
Model type Other
Context window 4K tokens
Release date 2024-12-02
Status Active
Knowledge cutoff notes

Cohere does not publish a specific knowledge cutoff for Rerank 3.5. The model evaluates supplied query and document content rather than relying on a disclosed conversational knowledge base.

Model notes

The canonical Cohere API model ID is rerank-v3.5. It ranks text documents or structured records serialized as YAML and returns document indexes with relevance scores rather than generated prose. Cohere documents a 4,096-token context length and automatic chunking for longer documents. The model is available through the Cohere API and has documented availability on Amazon Bedrock, Azure AI Foundry, and Oracle Cloud Infrastructure. Rerank pricing is based on searches rather than input and output token pricing; a current public self-serve per-search price was not verified. Cohere Model Vault lists Rerank 3.5 at $5.00 per hour for a medium-tier instance. Fine-tuning for rerank models is being retired according to Cohere's deprecation documentation.

Model guide

Cohere Rerank 3.5: Multilingual Reranking for Search and RAG

Cohere Rerank 3.5 is a specialized multilingual model that evaluates the relevance of retrieved documents to a user query and reorders them. It is designed for enterprise search, hybrid retrieval, and retrieval-augmented generation rather than conversation or content generation.

What is Cohere Rerank 3.5?

Cohere Rerank 3.5 is a specialized multilingual reranking model from Cohere. Its API model identifier is rerank-v3.5. Unlike a conversational language model, it does not produce an answer, summary, or piece of prose. Instead, it examines the relationship between a search query and a list of candidate documents, assigns relevance scores, and returns the candidates ordered by relevance.

Rerank 3.5 normally operates after an initial retrieval step. A keyword search engine, vector database, hybrid search system, or other index first finds potentially relevant documents. Rerank 3.5 then performs a more detailed comparison between the query and those candidates. The reordered results can be displayed directly or passed to a retrieval-augmented generation system as context for another model.

This makes Rerank 3.5 a search-quality component rather than a complete search product. It does not replace an index, document store, embedding model, database, or first-stage retrieval system.

How the model improves search results

Initial retrieval systems are usually optimized for speed. They may use keyword matching, vector similarity, or both to produce a relatively large candidate set. These methods are useful for finding possible matches, but the first results are not always the best semantic matches for the user's exact question.

Rerank 3.5 looks at each candidate in the context of the query and produces a relevance score normalized between 0 and 1. For example, an enterprise knowledge base might first retrieve 50 documents for a question about an expense policy. Rerank 3.5 can then identify which documents actually address reimbursement limits, rather than merely containing related words such as “expenses” or “travel.”

The model can rank plain text documents as well as structured information serialized as YAML strings. This allows records such as product listings, support tickets, policies, or project entries to be evaluated without converting every item into an unstructured paragraph.

Inputs, languages, and context limits

Cohere documents Rerank 3.5 as a multilingual model supporting more than 100 languages. The supported language set is broadly aligned with Cohere's multilingual embedding models, although actual ranking quality can vary between languages and use cases. Multilingual support is particularly relevant for organizations searching across international support content, policies, product information, or user-generated documents.

The model has a 4,096-token context length for each query-document evaluation. A token is a unit of text processed by the model and may represent part of a word, a complete word, or punctuation. The limit applies to the query and document content being evaluated, so applications working with long documents need to consider how the text is prepared.

Cohere's reranking service can automatically split longer documents into chunks. Automatic chunking is convenient, but applications that need predictable ranking behavior may choose their own chunk size and splitting strategy. Custom chunking also makes it easier to preserve headings, section boundaries, record identifiers, and other metadata that may matter when displaying results.

  • Text queries paired with candidate documents
  • Plain text or structured records represented as YAML strings
  • More than 100 supported languages according to Cohere's documentation
  • Up to 4,096 tokens for each query-document evaluation
  • Relevance scores normalized from 0 to 1
  • Support for requests containing up to 10,000 documents, subject to chunking and service limits

Where Rerank 3.5 fits in Cohere's lineup

Rerank 3.5 belongs to Cohere's retrieval model family. It complements, rather than replaces, models and systems used for embeddings, generation, or document processing. In a typical Cohere-based retrieval pipeline, an embedding or search system can identify candidate content, Rerank 3.5 can improve the ordering, and a separate generative model can write an answer using the highest-ranked passages.

The model was released on December 2, 2024, and is listed as active in the supplied model information. Its model family is “Rerank,” and its model type is classified as “other” rather than as a general-purpose text-generation model. Cohere documents the model through its own API and lists availability through services including Amazon Bedrock, Microsoft Azure AI Foundry, and Oracle Cloud Infrastructure.

Strengths and practical use cases

Rerank 3.5 is most useful when an application already retrieves plausible candidates but needs better ordering. Its main practical strengths are query-sensitive relevance evaluation, multilingual support, structured-data handling, and integration into enterprise retrieval workflows.

  • Enterprise search: Reorder documents from internal policies, technical documentation, human-resources content, or knowledge bases.
  • Retrieval-augmented generation: Select the most relevant passages before sending them to a separate answer-generating model, helping reduce irrelevant context.
  • Hybrid search: Improve results produced by combining keyword and vector retrieval.
  • Customer support: Rank help-center articles, internal troubleshooting records, or prior support cases against a customer's question.
  • Document discovery: Find relevant contracts, reports, emails, or project files across large enterprise repositories.
  • Structured search: Rank YAML-formatted product, finance, project, or operational records according to a natural-language query.
  • Multilingual retrieval: Search across content written in multiple languages when a single-language keyword strategy would be insufficient.

Because the model returns document references and scores instead of generated explanations, developers generally use the score to select, filter, or reorder results. The score is query-dependent: a score from one query should not automatically be treated as directly comparable with a score from another query without testing and calibration on representative application data.

Output, reasoning, coding, and tool support

Rerank 3.5 accepts text input and returns ranking information, including document indexes and relevance scores. It does not generate text, images, audio, video, embeddings, or other direct media outputs. It also does not function as a conversational reasoning or coding model.

There is no documented tool or function-calling capability for this model in the supplied specifications. Streaming is listed as unsupported, and the model is not described as supporting fine-tuning. These characteristics reflect its narrow ranking role: it evaluates supplied content rather than carrying out actions or producing a multi-step response.

Editorial capability classifications rate its reasoning and coding suitability low compared with general-purpose generation models. These are comparative editorial assessments, not provider-published benchmark scores. In practical terms, the model can perform relevance evaluation, but it should not be selected for code generation, general reasoning, summarization, or user-facing conversation.

Pricing and deployment

Cohere prices Rerank models by search rather than by generated output tokens. A current public self-serve per-search price for Rerank 3.5 was not verified in the supplied official pricing material, so a per-request or per-token figure should not be assumed.

Cohere's Model Vault documentation lists Rerank 3.5 at $5.00 per hour for a medium performance-tier instance. The same documentation indicates that longer-term monthly and annual rates may also be available, but the supplied research does not establish a single default recurring price for those alternatives. Model Vault pricing should therefore be treated separately from self-serve API search pricing.

The model is available through Cohere's managed services and is documented for deployment through Amazon Bedrock, Microsoft Azure AI Foundry, and Oracle Cloud Infrastructure. The appropriate deployment route depends on an organization's cloud, security, compliance, and infrastructure requirements.

Limitations and trade-offs

The most important limitation is that Rerank 3.5 is not a standalone search solution. It needs a candidate set from another retrieval system, so it adds a second processing stage and associated latency to the application. Sending too many candidates can also increase cost or processing time, making candidate-count selection an important design decision.

The 4,096-token evaluation context is smaller than the context windows documented for newer Rerank 4 models. Long documents may need to be split into chunks, and poor chunking can separate a relevant passage from the heading or metadata that gives it meaning. Teams should test chunking, candidate counts, and score thresholds using their own documents and queries.

Rerank 3.5 also cannot write an answer or explain why a result was selected in natural language. If the application needs conversational responses, summarization, code, or agent actions, a separate generative model is required. Similarly, if the main requirement is vector retrieval across a very large corpus, an embedding model and a suitable search index remain necessary.

When to choose Rerank 3.5

Choose Rerank 3.5 when your system already retrieves candidate documents and the main problem is ordering them more accurately according to a user's query. It is a good fit for multilingual enterprise search, hybrid retrieval, RAG pipelines, and structured records where semantic relevance matters more than simple keyword overlap.

Its narrower role can also be an advantage. Compared with using a general-purpose language model to judge search results, a dedicated reranker is designed specifically to score query-document relationships and return ranked candidates. The supplied editorial assessment also rates its speed and cost favorably, but these are comparative evaluations rather than guaranteed service-level measurements.

Another option may be more appropriate when the application needs generated answers, tool use, coding, image or audio processing, or a much larger per-document context. A newer reranking model may be preferable when its larger context window materially reduces the need for document chunking. Conversely, a simpler keyword or vector-only search system may be sufficient when very low latency or minimal processing cost matters more than the additional relevance stage.

Before deployment, evaluate Rerank 3.5 with real queries, languages, document formats, and failure cases. Pay particular attention to whether structured records are consistently serialized, whether long documents are chunked predictably, and whether relevance scores are being used as calibrated thresholds or merely as an ordering signal.


Answers to Frequently Asked Questions

How does Cohere Rerank 3.5 fit into a RAG pipeline?
A search engine, vector database, or hybrid retrieval system first finds candidate documents. Rerank 3.5 then evaluates and reorders those candidates by query relevance, after which the highest-ranked passages can be sent to a separate language model to generate the final response.
What is the context limit of Cohere Rerank 3.5?
Rerank 3.5 has a 4,096-token context limit for each query-document evaluation. Long documents may need to be split into chunks, either automatically by the service or through a custom chunking strategy.
How many languages and documents does Cohere Rerank 3.5 support?
Cohere documents Rerank 3.5 as supporting more than 100 languages. It can process requests containing up to 10,000 documents, subject to chunking and service limits, and supports plain-text documents as well as structured records serialized as YAML.
Does Cohere Rerank 3.5 generate answers or summaries?
No. Rerank 3.5 is a specialized ranking model, not a conversational or text-generation model. It returns document references, rankings, and relevance scores; a separate generative model is required to write answers, summaries, or explanations.
What is Cohere Rerank 3.5 used for?
Cohere Rerank 3.5 is used to reorder candidate documents according to their relevance to a search query. It is commonly used in enterprise search, hybrid search, multilingual retrieval, structured-data search, and retrieval-augmented generation (RAG) pipelines.


Sources 9
Provider

About Cohere