Multilingual generation, translation, summarization, customer support, global communication, and multilingual research
Type
General Purpose
Context
128K
Reasoning
6/10
Speed
4/10
Fine-tuning
Input
$0.50 per 1 million tokens
Output
$1.50 per 1 million tokens
View model
→
Multilingual image understanding, OCR, image captioning, visual question answering, image-based translation, visual reasoning and research deployments using open weights.
Type
Multimodal
Context
16K
Reasoning
7/10
Speed
5/10
Multimodal
Image input
Streaming
View model
→
Multilingual audio transcription, enterprise speech archives, meeting and interview transcription, call-center audio, and low-latency ASR workflows
Type
Speech Recognition
Reasoning
2/10
Speed
8/10
Audio input
Status
Live; open-source research release
Input
No per-input token price published. API experimentation is free subject to rate limits; Model Vault uses hourly instance-based pricing.
Output
No per-output token price published. API experimentation is free subject to rate limits; Model Vault uses hourly instance-based pricing.
View model
→
Arabic speech transcription, multilingual Arabic-English audio, regional dialects, code-switched speech, call-center audio and high-throughput ASR workloads.
Type
Other
Reasoning
1/10
Speed
8/10
Audio input
Input
Free through Cohere API for experimentation subject to rate limits; Model Vault deployment is priced per hour-instance and requires contacting Cohere.
View model
→
Enterprise RAG, long-context document analysis, multilingual applications, tool use, agentic workflows, financial text processing and structured text generation
Type
General Purpose
Context
256K
Reasoning
8/10
Speed
8/10
Tool use
Streaming
Input
$2.50 per 1 million tokens
Output
$10.00 per 1 million tokens
View model
→
Enterprise agents, multimodal document and image analysis, multilingual workflows, reasoning-intensive automation, retrieval-augmented generation, and tool-using applications.
Type
Multimodal
Context
128K
Reasoning
9/10
Speed
9/10
Multimodal
Image input
Tool use
Input
Free until applicable rate limits; production pricing is not publicly listed and may require contacting Cohere.
Output
Free until applicable rate limits; production pricing is not publicly listed and may require contacting Cohere.
View model
→
Complex enterprise agents, tool use, retrieval-augmented generation, multilingual reasoning, long-context analysis, and workflow automation.
Type
Reasoning
Context
256K
Reasoning
8/10
Speed
5/10
Tool use
Web search
Streaming
Input
Free until applicable API rate limits; private Model Vault deployment is billed per instance-hour rather than per token. Standard Vault pricing lists $48/hour for L and $57.50/hour for XL performance tiers.
Output
Free until applicable API rate limits on Cohere-hosted trial and production access; Model Vault pricing is instance-hour based rather than output-token based.
View model
→
High-quality multilingual text translation, enterprise document translation, and privacy-sensitive translation workflows.
Type
Specialized
Context
8K
Reasoning
4/10
Speed
7/10
Tool use
Streaming
Input
Free until applicable rate limits are reached; production access requires contacting Cohere sales.
Output
Free until applicable rate limits are reached; production access requires contacting Cohere sales.
View model
→
Enterprise document intelligence, OCR, chart and table analysis, visual question answering, and multilingual image understanding
Type
Multimodal
Context
128K
Reasoning
6/10
Speed
7/10
Multimodal
Image input
Streaming
Input
Free until applicable rate limits; production access requires contacting Cohere
Output
Free until applicable rate limits; production access requires contacting Cohere
View model
→
Low-cost enterprise RAG, multilingual document workflows, long-context chat, structured extraction, and tool-using agents
Type
General Purpose
Context
128K
Reasoning
6/10
Speed
8/10
Tool use
Fine-tuning
Streaming
Input
$0.15 per 1M input tokens
Output
$0.60 per 1M output tokens
View model
→
Cost-sensitive RAG, enterprise chat, tool use, coding assistance, and fast multi-step agents
Type
Lightweight
Context
128K
Reasoning
7/10
Speed
9/10
Tool use
Web search
Streaming
Input
$0.0375 per 1 million tokens
Output
$0.15 per 1 million tokens
View model
→
Complex enterprise RAG, long-context document analysis, multilingual assistants, citations, structured data tasks and multi-step tool-use agents.
Type
General Purpose
Context
128K
Reasoning
7/10
Speed
7/10
Tool use
Streaming
Input
$2.50 per 1 million tokens
Output
$10 per 1 million tokens
View model
→
Multilingual semantic search, multimodal RAG, PDF and document retrieval, image-to-text retrieval, classification, clustering, and enterprise vector indexing.
Type
Multimodal
Context
128K
Reasoning
1/10
Speed
8/10
Multimodal
Image input
Status
Generally available
Input
$0.12 per 1 million text input tokens; $0.47 per 1 million image tokens
View model
→
Fast, storage-efficient English semantic search, retrieval, classification, clustering, and large-scale embedding workloads
Type
Lightweight
Context
512
Speed
9/10
Multimodal
Image input
Status
Current and available
View model
→
English semantic search, retrieval-augmented generation, vector indexing, classification, clustering, and similarity matching
Type
Other
Context
512
Speed
7/10
Multimodal
Image input
Input
$0.10 per 1 million input tokens
Output
Not applicable; the model returns embeddings rather than billed generated text
View model
→
Fast multilingual semantic search, cross-lingual retrieval, RAG, clustering, classification features, and compact vector indexes
Type
Lightweight
Context
512
Speed
9/10
Multimodal
Image input
View model
→
Multilingual semantic search, cross-lingual retrieval, RAG indexing, classification, clustering, and image-text similarity
Type
Other
Context
512
Reasoning
1/10
Speed
8/10
Multimodal
Image input
Input
$0.10 per 1M input tokens
View model
→
Agentic software engineering, repository-level code changes, terminal-based coding agents, code review, local inference, and private deployment.
Type
Coding
Context
256K
Reasoning
7/10
Speed
8/10
Multimodal
Image input
Tool use
Input
Free until applicable Cohere API rate limits are reached; self-hosted deployment has infrastructure costs rather than Cohere per-token pricing.
Output
Free until applicable Cohere API rate limits are reached; self-hosted deployment has infrastructure costs rather than Cohere per-token pricing.
View model
→
Enterprise machine translation, multilingual documentation, localization, internal communications, safety procedures, and private or self-hosted translation workflows
Type
Other
Context
16K
Reasoning
6/10
Speed
6/10
Tool use
Streaming
Input
Cohere API: free until applicable rate limits are reached; Standard Model Vault: $17.50 per instance-hour for XL deployment
Output
No separate token-based output price published; Cohere API is free until applicable rate limits are reached
View model
→
Compact multilingual image understanding, OCR, documents, charts, visual question answering, and specialized fine-tuning
Type
Multimodal
Context
128K
Reasoning
5/10
Speed
8/10
Multimodal
Image input
Fine-tuning
Status
Current open-weight release
View model
→
High-volume enterprise document parsing, table and form extraction, search indexing, RAG ingestion, and document context for AI agents
Type
Multimodal
Context
8K
Reasoning
2/10
Speed
8/10
Multimodal
Image input
Status
Live; generally available
Input
$1.50 per 1,000 pages through the Cohere API
View model
→
Multilingual semantic reranking, enterprise search, hybrid retrieval, and RAG pipelines
Type
Other
Context
4K
Reasoning
3/10
Speed
8/10
View model
→
English semantic reranking for enterprise search, hybrid retrieval, RAG pipelines, FAQs, knowledge bases, documents, code retrieval, and semi-structured records
Type
Other
Context
4K
Reasoning
2/10
Speed
8/10
Status
Active and currently listed by Cohere; older Rerank 3.0 English model with newer Rerank alternatives available
Input
$2.00 per 1,000 search units; one search unit covers one query with up to 100 documents, subject to chunking
View model
→
Multilingual semantic reranking for enterprise search, hybrid retrieval, cross-language search, and RAG pipelines
Type
Other
Context
4K
Reasoning
5/10
Speed
7/10
Status
Current older-generation model; superseded by newer Rerank model generations
View model
→
Low-latency multilingual search reranking, high-throughput retrieval, enterprise search, hybrid search and RAG pipelines
Type
Lightweight
Context
33K
Reasoning
1/10
Speed
9/10
Status
Current and available
Input
Usage-based search-unit pricing; one search unit is one query with up to 100 documents to be ranked
View model
→
High-quality multilingual reranking for enterprise search, RAG pipelines, semantic retrieval, and semi-structured document ranking.
Type
Other
Context
33K
Reasoning
4/10
Speed
7/10
Status
Generally available
Input
$2.50 per 1,000 search units; one search unit is one query against up to 100 documents. Dedicated Model Vault pricing is separate.
Output
Not applicable; the endpoint returns relevance scores and ranking results rather than generated tokens.
View model
→
Multilingual translation, conversation, summarization, and text generation focused on African and West Asian languages; local and edge deployment
Type
Lightweight
Context
8K
Reasoning
3/10
Speed
8/10
Fine-tuning
View model
→
South Asian multilingual conversation, translation, target-language generation, local inference, and edge or on-device applications
Type
Lightweight
Context
8K
Reasoning
3/10
Speed
8/10
View model
→
Multilingual translation, cross-lingual text generation, localized assistants, education, research, and efficient local or edge deployment
Type
Lightweight
Context
8K
Reasoning
3/10
Speed
8/10
Fine-tuning
View model
→
Efficient multilingual translation, text generation, localization, language learning, and local or edge deployment for European and Asia-Pacific languages
Type
Lightweight
Context
8K
Reasoning
3/10
Speed
8/10
Fine-tuning
Status
Live; available through the Cohere Chat API and as an open-weight model
View model
→