What is Cohere Rerank 3.5?
Cohere Rerank 3.5 is a specialized multilingual reranking model from Cohere. Its API model identifier is rerank-v3.5. Unlike a conversational language model, it does not produce an answer, summary, or piece of prose. Instead, it examines the relationship between a search query and a list of candidate documents, assigns relevance scores, and returns the candidates ordered by relevance.
Rerank 3.5 normally operates after an initial retrieval step. A keyword search engine, vector database, hybrid search system, or other index first finds potentially relevant documents. Rerank 3.5 then performs a more detailed comparison between the query and those candidates. The reordered results can be displayed directly or passed to a retrieval-augmented generation system as context for another model.
This makes Rerank 3.5 a search-quality component rather than a complete search product. It does not replace an index, document store, embedding model, database, or first-stage retrieval system.
How the model improves search results
Initial retrieval systems are usually optimized for speed. They may use keyword matching, vector similarity, or both to produce a relatively large candidate set. These methods are useful for finding possible matches, but the first results are not always the best semantic matches for the user's exact question.
Rerank 3.5 looks at each candidate in the context of the query and produces a relevance score normalized between 0 and 1. For example, an enterprise knowledge base might first retrieve 50 documents for a question about an expense policy. Rerank 3.5 can then identify which documents actually address reimbursement limits, rather than merely containing related words such as “expenses” or “travel.”
The model can rank plain text documents as well as structured information serialized as YAML strings. This allows records such as product listings, support tickets, policies, or project entries to be evaluated without converting every item into an unstructured paragraph.
Inputs, languages, and context limits
Cohere documents Rerank 3.5 as a multilingual model supporting more than 100 languages. The supported language set is broadly aligned with Cohere's multilingual embedding models, although actual ranking quality can vary between languages and use cases. Multilingual support is particularly relevant for organizations searching across international support content, policies, product information, or user-generated documents.
The model has a 4,096-token context length for each query-document evaluation. A token is a unit of text processed by the model and may represent part of a word, a complete word, or punctuation. The limit applies to the query and document content being evaluated, so applications working with long documents need to consider how the text is prepared.
Cohere's reranking service can automatically split longer documents into chunks. Automatic chunking is convenient, but applications that need predictable ranking behavior may choose their own chunk size and splitting strategy. Custom chunking also makes it easier to preserve headings, section boundaries, record identifiers, and other metadata that may matter when displaying results.
- Text queries paired with candidate documents
- Plain text or structured records represented as YAML strings
- More than 100 supported languages according to Cohere's documentation
- Up to 4,096 tokens for each query-document evaluation
- Relevance scores normalized from 0 to 1
- Support for requests containing up to 10,000 documents, subject to chunking and service limits
Where Rerank 3.5 fits in Cohere's lineup
Rerank 3.5 belongs to Cohere's retrieval model family. It complements, rather than replaces, models and systems used for embeddings, generation, or document processing. In a typical Cohere-based retrieval pipeline, an embedding or search system can identify candidate content, Rerank 3.5 can improve the ordering, and a separate generative model can write an answer using the highest-ranked passages.
The model was released on December 2, 2024, and is listed as active in the supplied model information. Its model family is “Rerank,” and its model type is classified as “other” rather than as a general-purpose text-generation model. Cohere documents the model through its own API and lists availability through services including Amazon Bedrock, Microsoft Azure AI Foundry, and Oracle Cloud Infrastructure.
Strengths and practical use cases
Rerank 3.5 is most useful when an application already retrieves plausible candidates but needs better ordering. Its main practical strengths are query-sensitive relevance evaluation, multilingual support, structured-data handling, and integration into enterprise retrieval workflows.
- Enterprise search: Reorder documents from internal policies, technical documentation, human-resources content, or knowledge bases.
- Retrieval-augmented generation: Select the most relevant passages before sending them to a separate answer-generating model, helping reduce irrelevant context.
- Hybrid search: Improve results produced by combining keyword and vector retrieval.
- Customer support: Rank help-center articles, internal troubleshooting records, or prior support cases against a customer's question.
- Document discovery: Find relevant contracts, reports, emails, or project files across large enterprise repositories.
- Structured search: Rank YAML-formatted product, finance, project, or operational records according to a natural-language query.
- Multilingual retrieval: Search across content written in multiple languages when a single-language keyword strategy would be insufficient.
Because the model returns document references and scores instead of generated explanations, developers generally use the score to select, filter, or reorder results. The score is query-dependent: a score from one query should not automatically be treated as directly comparable with a score from another query without testing and calibration on representative application data.
Output, reasoning, coding, and tool support
Rerank 3.5 accepts text input and returns ranking information, including document indexes and relevance scores. It does not generate text, images, audio, video, embeddings, or other direct media outputs. It also does not function as a conversational reasoning or coding model.
There is no documented tool or function-calling capability for this model in the supplied specifications. Streaming is listed as unsupported, and the model is not described as supporting fine-tuning. These characteristics reflect its narrow ranking role: it evaluates supplied content rather than carrying out actions or producing a multi-step response.
Editorial capability classifications rate its reasoning and coding suitability low compared with general-purpose generation models. These are comparative editorial assessments, not provider-published benchmark scores. In practical terms, the model can perform relevance evaluation, but it should not be selected for code generation, general reasoning, summarization, or user-facing conversation.
Pricing and deployment
Cohere prices Rerank models by search rather than by generated output tokens. A current public self-serve per-search price for Rerank 3.5 was not verified in the supplied official pricing material, so a per-request or per-token figure should not be assumed.
Cohere's Model Vault documentation lists Rerank 3.5 at $5.00 per hour for a medium performance-tier instance. The same documentation indicates that longer-term monthly and annual rates may also be available, but the supplied research does not establish a single default recurring price for those alternatives. Model Vault pricing should therefore be treated separately from self-serve API search pricing.
The model is available through Cohere's managed services and is documented for deployment through Amazon Bedrock, Microsoft Azure AI Foundry, and Oracle Cloud Infrastructure. The appropriate deployment route depends on an organization's cloud, security, compliance, and infrastructure requirements.
Limitations and trade-offs
The most important limitation is that Rerank 3.5 is not a standalone search solution. It needs a candidate set from another retrieval system, so it adds a second processing stage and associated latency to the application. Sending too many candidates can also increase cost or processing time, making candidate-count selection an important design decision.
The 4,096-token evaluation context is smaller than the context windows documented for newer Rerank 4 models. Long documents may need to be split into chunks, and poor chunking can separate a relevant passage from the heading or metadata that gives it meaning. Teams should test chunking, candidate counts, and score thresholds using their own documents and queries.
Rerank 3.5 also cannot write an answer or explain why a result was selected in natural language. If the application needs conversational responses, summarization, code, or agent actions, a separate generative model is required. Similarly, if the main requirement is vector retrieval across a very large corpus, an embedding model and a suitable search index remain necessary.
When to choose Rerank 3.5
Choose Rerank 3.5 when your system already retrieves candidate documents and the main problem is ordering them more accurately according to a user's query. It is a good fit for multilingual enterprise search, hybrid retrieval, RAG pipelines, and structured records where semantic relevance matters more than simple keyword overlap.
Its narrower role can also be an advantage. Compared with using a general-purpose language model to judge search results, a dedicated reranker is designed specifically to score query-document relationships and return ranked candidates. The supplied editorial assessment also rates its speed and cost favorably, but these are comparative evaluations rather than guaranteed service-level measurements.
Another option may be more appropriate when the application needs generated answers, tool use, coding, image or audio processing, or a much larger per-document context. A newer reranking model may be preferable when its larger context window materially reduces the need for document chunking. Conversely, a simpler keyword or vector-only search system may be sufficient when very low latency or minimal processing cost matters more than the additional relevance stage.
Before deployment, evaluate Rerank 3.5 with real queries, languages, document formats, and failure cases. Pay particular attention to whether structured records are consistently serialized, whether long documents are chunked predictably, and whether relevance scores are being used as calibrated thresholds or merely as an ordering signal.

