/v2/chat, then add Embed, Rerank, Parse or batch APIs as the application requires. This guide explains how access works, how to make a first request, and what to consider before moving from a trial key to production.
What is the Cohere API and when should you use it?
The Cohere API is a developer platform for adding language generation, semantic retrieval and document intelligence to applications. Its current generative interface is the Chat API v2, which accepts an ordered messages array containing system, user, assistant and tool messages. It is intended for applications rather than casual consumer chat use.
The platform is especially relevant when an application needs enterprise search, retrieval-augmented generation (RAG), multilingual responses, document analysis, tool-driven workflows or deployment controls such as private infrastructure, Model Vault or cloud-provider integrations. RAG means retrieving relevant information from a connected data source and giving that information to the model as context before it generates an answer.
Cohere is a less natural fit for a consumer chatbot, image or video generation product, or a general-purpose personal assistant. Its APIs provide building blocks, but your application remains responsible for user interfaces, data pipelines, tool execution, permissions and most workflow orchestration.
How do you get access and an API key?
Create a Cohere account and obtain a trial API key through the Cohere developer dashboard. The key authenticates requests with an HTTP Authorization: Bearer header. Store it in an environment variable rather than placing it directly in source code or a browser application.
Trial keys are intended for development and are rate-limited. The supplied platform information describes trial usage as limited and not suitable for production or commercial use. Production access requires upgrading through Cohere's account and billing workflow.
export COHERE_API_KEY="your-api-key"The dashboard also provides a playground for trying requests interactively: dashboard.cohere.com/playground. Use it to inspect model behavior and request formats before writing application code, but test authentication, error handling and rate-limit behavior in your own integration as well.
Which Cohere API and model should you choose?
Choose the API according to the job rather than treating Chat as the only interface:
- Chat API v2: conversational generation, RAG, citations, streaming, tool calling, structured output and compatible image inputs.
- Embed API: converts text or supported images into vectors that can be used for semantic search, classification and retrieval.
- Rerank API: reorders search results or retrieved documents by relevance before they are passed to a generative model.
- Parse API: converts supported documents and images into markdown or structured blocks for downstream processing.
- Datasets and Embed Jobs: handle file-backed asynchronous embedding workflows for larger corpora.
- Batch API: processes uploaded request datasets asynchronously for supported workloads.
Model compatibility varies by capability. Select an active model that supports the features you need, such as tools, structured output or image input, and verify the current model documentation before deployment. Legacy generation endpoints and older model aliases should not be used for new integrations.
How do you make a first request?
The REST base URL is https://api.cohere.com. The following request uses Chat API v2. Replace the example model identifier with an active model available to your account if necessary.
curl --fail-with-body --silent --show-error https://api.cohere.com/v2/chat
-H "Authorization: Bearer ${COHERE_API_KEY}"
-H "Content-Type: application/json"
-H "Accept: application/json"
-d '{
"model": "command-a-plus-05-2026",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain retrieval-augmented generation in two sentences."}
]
}'The same request can be made with Cohere's current Python SDK. Install the package with pip install -U cohere and use the v2 client.
import os
import cohere
client = cohere.ClientV2(api_key=os.environ["COHERE_API_KEY"])
response = client.chat(
model="command-a-plus-05-2026",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain semantic search in two sentences."}
],
)
text = "".join(
part.text for part in response.message.content
if part.type == "text"
)
print(text)How should you read the response?
A Chat API v2 response contains a message with content parts. Text content is represented as a content part whose type is text. Applications should collect the text parts rather than assuming the response is always one plain string, because a response can also contain tool calls, citations or other event types depending on the request.
For production code, handle unsuccessful HTTP responses, empty or unexpected content, rate-limit responses and model-specific errors. Preserve usage metadata when available so that your application can measure token consumption and investigate cost or latency changes.
How does Cohere API pricing work?
Cohere uses usage-based pricing that varies by model and endpoint. Generative models are generally priced by input and output tokens. Reranking is priced by search volume or applicable units, while embeddings are priced by embedded tokens or workload units. Dedicated or managed deployment options have separate pricing.
Trial keys are free but rate-limited. The supplied rate-limit information describes trial keys as generally limited to 1,000 API calls per month and commonly 20 requests per minute for Chat models. Production limits vary by model and endpoint; examples in the supplied documentation include as many as 500 requests per minute for several active Chat models, 2,000 inputs per minute for Embed and 1,000 requests per minute for Rerank. Some newer models require contacting sales for production limits.
Check Cohere's current pricing and rate-limit documentation before estimating costs. Do not assume that a limit or price for one model applies to another.
What can developers build with the API?
Streaming responses
Chat responses can be streamed as server-sent events. Instead of waiting for the complete answer, the application receives events as content is generated. This is useful for user interfaces where displaying the first words quickly matters. Stream events may include message starts, content deltas, tool calls, citations and completion events.
stream = client.chat_stream(
model="command-a-plus-05-2026",
messages=[
{"role": "user", "content": "Explain embeddings in three short steps."}
],
)
for event in stream:
if event.type == "content-delta" and event.delta and event.delta.message:
print(event.delta.message, end="", flush=True)
print()Embeddings, reranking and RAG
A typical RAG pipeline uses Embed to represent documents and user queries as vectors, a search system to find candidate passages, and Rerank to improve the ordering of those candidates. The selected passages are then supplied to Chat as context. This separates retrieval from generation and lets the application control which enterprise information is shown to the model.
Cohere also supports citations and retrieval documents in Chat workflows. Your application should still enforce access permissions before sending retrieved content to the model.
Tool and function calling
Tools let the model request an action such as querying an internal system or looking up an order. The model does not execute the tool itself. Your application receives a tool call, validates the requested arguments, performs the action, and sends the result back in a follow-up message.
Tool definitions use a JSON Schema-like description. Cohere supports multi-step and parallel tool patterns, but the application remains responsible for authorization, validation, retries, timeouts and preventing unsafe actions. The experimental strict_tools option can enforce tool names, required parameters and parameter types on compatible Chat API v2 workflows.
Structured outputs
Chat API v2 supports the response_format parameter for JSON mode and JSON Schema-style output on compatible models. For example:
response = client.chat(
model="command-a-plus-05-2026",
messages=[
{"role": "user", "content": "Describe semantic search with fields name and benefit."}
],
response_format={"type": "json_object"},
)
print(response.message.content[0].text)JSON mode does not remove the need to validate the result. Model compatibility varies, and structured output should be treated as a response-format feature rather than a separate JSON-only API.
Images, documents and files
Compatible vision models accept images through HTTP URLs or base64 data URLs. Supported image formats include PNG, JPEG, WEBP and non-animated GIF, subject to documented image-count and request-size limits.
Cohere also supports CSV and JSONL dataset uploads for embedding jobs and provides document and image parsing through the Parse API. These uploads are task-specific; they are not a universal, persistent file-search store. Dataset files are automatically deleted after the documented retention period of 30 days unless deleted sooner.
Which SDKs and developer tools are available?
Cohere officially supports Python, TypeScript, Java and Go SDKs. REST and HTTP access can be used from other languages, including PHP. The current Python client uses cohere.ClientV2; the current TypeScript package uses CohereClientV2 from cohere-ai.
Use the official API reference, model documentation, SDK repositories, playground, cookbooks and deprecation notices together. The playground is useful for exploration, while automated tests should verify the exact request and response structures your application depends on.
What limits and production issues matter?
- Rate limits: Limits differ between trial and production keys and between models and endpoints. Implement backoff and handle rate-limit responses.
- Latency: There is no single universal latency guarantee. Model choice, prompt size, output length, image detail, tools, queue priority and deployment affect response time. Streaming can reduce perceived waiting time.
- Model compatibility: Tools, structured output and image input are not necessarily available on every model. Test each capability with the selected model.
- Data retention: Retention depends on the product and deployment. Model Vault can support Zero Data Retention configurations, while private and partner-cloud deployments can provide stronger customer-controlled data boundaries.
- Security: Keep API keys server-side, validate tool arguments, filter retrieved content and apply application-level authorization before exposing enterprise data or actions.
- Deprecations: Legacy endpoints such as
/v1/generate,/v1/summarizeand/v1/classify, along with related legacy features, should not be the basis of new applications. - Fine-tuning: Cohere's deprecation documentation states that the listed platform fine-tuning capabilities were retired effective September 15, 2025. Do not plan a new integration around historical fine-tuning endpoints.
Is Cohere API a good choice for your project?
Cohere is a strong candidate when an organization needs multilingual generation, semantic retrieval, reranking, document processing, agent workflows or deployment options that include private, dedicated, cloud or air-gapped environments. The separation between Chat, Embed, Rerank and Parse is useful for building controlled enterprise search and RAG systems.
It may be a poor choice if the main requirement is a free consumer chatbot, first-party mobile or desktop applications, image or video generation, or a fully managed personal productivity ecosystem. Enterprise pricing is often custom, advanced workflows require application-side engineering, and capabilities can vary by model and deployment. Evaluate the complete workflow—including data governance, retrieval quality, operational limits and tool safety—rather than testing text generation alone.
