Developer platform

Cohere API

Developer overview for Cohere, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Chat API v2
SDK support Official SDKs for Python, TypeScript, Java and Go. REST/HTTP access is available for other languages, including PHP.
Rate limits Trial keys are generally limited to 1,000 API calls per month and commonly 20 requests per minute for Chat models. Published production limits vary by model and endpoint; examples include up to 500 requests per minute for several active Chat models, 2,000
Platform

API overview

Endpoints

API access

Base URL https://api.cohere.com
Primary API Chat API v2
Pricing

API pricing

Pricing model Usage-based pricing by model and endpoint; token-based for generative models, search-unit pricing for reranking, token-based pricing for embeddings, and separate deployment pricing for managed dedicated infrastructure.

Trial API keys are free but rate-limited. Production usage is metered. Current model pricing varies; Command A+ is listed as free until applicable rate limits are reached, while other models and endpoints have published per-token or per-unit prices.

Developer experience

SDKs & usability

SDKs Official SDKs for Python, TypeScript, Java and Go. REST/HTTP access is available for other languages, including PHP.
Ease of use Good. The REST API is conventional, the v2 Chat API uses a clear messages structure, and official SDKs provide typed clients and helpers. Advanced RAG and tool workflows require more application-side orchestration.
Documentation Good to very good. Cohere provides current guides, API references, SDK repositories, model pages, migration guidance, rate-limit documentation, cookbooks and deprecation notices.
Latency No single universal latency guarantee is published for the public API. Latency depends on model, prompt size, output length, image detail, tools, queue priority and deployment. The API supports a priority parameter and streaming to reduce perceived time t
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The current recommended generative endpoint is POST /v2/chat, with POST /v2/chat_stream or SDK chat_stream methods for streaming. Chat v2 supports system, user, assistant and tool messages, retrieval documents, citations, tool/function calling, structured JSON output and compatible image inputs. Tool use can support multi-step agent workflows, but the application remains responsible for executing tools and returning tool results. Structured JSON output is model-dependent; strict_tools is currently documented as experimental and is available only in Chat API v2. Image input is available on compatible vision models, including Command A Vision and newer multimodal models. Cohere provides dataset uploads for CSV and JSONL files used by Embed Jobs, but this is not equivalent to a universal persistent file-search API. Fine-tuning should not be treated as a current supported platform capability because the deprecation documentation states that Cohere retired the listed fine-tuning capabilities effective September 15, 2025. Legacy /v1/generate, /v1/summarize, /v1/classify, managed connectors and related parameters are deprecated; new applications should use Chat v2 and current specialized APIs. Cohere also offers Model Vault, private deployments and cloud deployments for enterprise privacy, isolation and deployment-control requirements. The provider offers agentic capabilities through tool use, agent-building workflows and enterprise products such as North.

Examples

API examples

curl --fail-with-body --silent --show-error https://api.cohere.com/v2/chat \
  -H "Authorization: Bearer ${COHERE_API_KEY}" \
  -H "Content-Type: application/json" \
  -H "Accept: application/json" \
  -d '{
    "model": "command-a-plus-05-2026",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain retrieval-augmented generation in two sentences."}
    ]
  }' | jq -r '.message.content[] | select(.type == "text") | .text'

curl --fail-with-body --silent --show-error https://api.cohere.com/v2/chat \
  -H "Authorization: Bearer ${COHERE_API_KEY}" \
  -H "Content-Type: application/json" \
  -H "Accept: text/event-stream" \
  -d '{
    "model": "command-a-plus-05-2026",
    "messages": [{"role": "user", "content": "Stream a short explanation of semantic search."}],
    "stream": true
  }'
# Install: pip install -U cohere
import json
import os
import cohere

client = cohere.ClientV2(api_key=os.environ["COHERE_API_KEY"])

try:
    response = client.chat(
        model="command-a-plus-05-2026",
        messages=[
            {"role": "system", "content": "You are a concise technical assistant."},
            {"role": "user", "content": "Return a JSON object describing semantic search with fields name and benefit."}
        ],
        response_format={"type": "json_object"},
    )
    print(response.message.content[0].text)

    stream = client.chat_stream(
        model="command-a-plus-05-2026",
        messages=[
            {"role": "user", "content": "Explain embeddings in three short steps."}
        ],
    )
    for event in stream:
        if event.type == "content-delta" and event.delta and event.delta.message:
            print(event.delta.message, end="", flush=True)
    print()
except Exception as exc:
    print(f"Cohere request failed: {exc}")
    raise
// Install: npm install cohere-ai
import { CohereClientV2, CohereError } from "cohere-ai";

const client = new CohereClientV2({
  token: process.env.COHERE_API_KEY,
});

try {
  const response = await client.chat({
    model: "command-a-plus-05-2026",
    messages: [
      { role: "system", content: "You are a concise technical assistant." },
      { role: "user", content: "Give a two-sentence explanation of RAG." },
    ],
  });

  const text = response.message?.content
    ?.filter((part) => part.type === "text")
    .map((part) => part.text)
    .join("");
  console.log(text);

  const stream = await client.chatStream({
    model: "command-a-plus-05-2026",
    messages: [
      { role: "user", content: "Stream three benefits of semantic search." },
    ],
  });

  for await (const event of stream) {
    if (event.type === "content-delta" && event.delta?.message) {
      process.stdout.write(event.delta.message);
    }
  }
  process.stdout.write("\n");
} catch (error) {
  if (error instanceof CohereError) {
    console.error(`Cohere error ${error.statusCode}: ${error.message}`);
  } else {
    console.error(error);
  }
  process.exitCode = 1;
}
<?php
$apiKey = getenv('COHERE_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('COHERE_API_KEY is not set');
}

$url = 'https://api.cohere.com/v2/chat';
$payload = [
    'model' => 'command-a-plus-05-2026',
    'messages' => [
        ['role' => 'system', 'content' => 'You are a concise technical assistant.'],
        ['role' => 'user', 'content' => 'Explain semantic search in two sentences.'],
    ],
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_TIMEOUT => 60,
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('cURL error: ' . $error);
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);

if ($status < 200 || $status >= 300) {
    $message = $data['message'] ?? $body;
    throw new RuntimeException('Cohere API error ' . $status . ': ' . $message);
}

$text = '';
foreach (($data['message']['content'] ?? []) as $part) {
    if (($part['type'] ?? null) === 'text') {
        $text .= $part['text'] ?? '';
    }
}

echo $text . PHP_EOL;
Policies

Data & usage

Data training

For Cohere-hosted environments, de-identified data from model use may be used in limited circumstances where permitted by user controls and applicable terms. Cohere states that customers using private or cloud deployments can keep data under their control and that, in most such environments, Cohere does not have access to submitted inputs. Organizations should review the current Cohere data statement, terms and deployment-specific controls before sending sensitive data.

Data retention

Retention depends on the product and deployment. Cohere-hosted API processing should be evaluated under the current data statement and account controls. Uploaded Datasets used for Embed Jobs are automatically deleted 30 days after creation unless deleted sooner. Model Vault supports Zero Data Retention configurations in which prompts and responses are not retained after inference. Private and partner cloud deployments may provide customer-controlled retention and data boundaries.

Rate limits

Trial keys are generally limited to 1,000 API calls per month and commonly 20 requests per minute for Chat models. Published production limits vary by model and endpoint; examples include up to 500 requests per minute for several active Chat models, 2,000

Developer guide

Cohere API: A Practical Guide to Chat, Search, RAG and Agents

Cohere provides an enterprise-focused developer platform built around the Chat API v2 and specialized APIs for embeddings, reranking, document parsing, batch processing and retrieval-augmented generation. Developers can authenticate with API keys, use official SDKs or REST, stream responses, call external tools, request structured JSON, process compatible image inputs and choose hosted, private or dedicated deployment options.
Cohere's API is designed for developers building enterprise search, multilingual assistants, document workflows, retrieval-augmented generation and agent applications. New projects should generally start with the Chat API v2 at /v2/chat, then add Embed, Rerank, Parse or batch APIs as the application requires. This guide explains how access works, how to make a first request, and what to consider before moving from a trial key to production.
Cohere API is an enterprise-oriented developer platform centered on Chat API v2, with additional APIs for embeddings, reranking, parsing and batch workloads. It supports official SDKs, streaming, multi-turn conversations, tool calling, structured JSON output, compatible image inputs and private or dedicated deployment options.

What is the Cohere API and when should you use it?

The Cohere API is a developer platform for adding language generation, semantic retrieval and document intelligence to applications. Its current generative interface is the Chat API v2, which accepts an ordered messages array containing system, user, assistant and tool messages. It is intended for applications rather than casual consumer chat use.

The platform is especially relevant when an application needs enterprise search, retrieval-augmented generation (RAG), multilingual responses, document analysis, tool-driven workflows or deployment controls such as private infrastructure, Model Vault or cloud-provider integrations. RAG means retrieving relevant information from a connected data source and giving that information to the model as context before it generates an answer.

Cohere is a less natural fit for a consumer chatbot, image or video generation product, or a general-purpose personal assistant. Its APIs provide building blocks, but your application remains responsible for user interfaces, data pipelines, tool execution, permissions and most workflow orchestration.

How do you get access and an API key?

Create a Cohere account and obtain a trial API key through the Cohere developer dashboard. The key authenticates requests with an HTTP Authorization: Bearer header. Store it in an environment variable rather than placing it directly in source code or a browser application.

Trial keys are intended for development and are rate-limited. The supplied platform information describes trial usage as limited and not suitable for production or commercial use. Production access requires upgrading through Cohere's account and billing workflow.

export COHERE_API_KEY="your-api-key"

The dashboard also provides a playground for trying requests interactively: dashboard.cohere.com/playground. Use it to inspect model behavior and request formats before writing application code, but test authentication, error handling and rate-limit behavior in your own integration as well.

Which Cohere API and model should you choose?

Choose the API according to the job rather than treating Chat as the only interface:

  • Chat API v2: conversational generation, RAG, citations, streaming, tool calling, structured output and compatible image inputs.
  • Embed API: converts text or supported images into vectors that can be used for semantic search, classification and retrieval.
  • Rerank API: reorders search results or retrieved documents by relevance before they are passed to a generative model.
  • Parse API: converts supported documents and images into markdown or structured blocks for downstream processing.
  • Datasets and Embed Jobs: handle file-backed asynchronous embedding workflows for larger corpora.
  • Batch API: processes uploaded request datasets asynchronously for supported workloads.

Model compatibility varies by capability. Select an active model that supports the features you need, such as tools, structured output or image input, and verify the current model documentation before deployment. Legacy generation endpoints and older model aliases should not be used for new integrations.

How do you make a first request?

The REST base URL is https://api.cohere.com. The following request uses Chat API v2. Replace the example model identifier with an active model available to your account if necessary.

curl --fail-with-body --silent --show-error https://api.cohere.com/v2/chat 
  -H "Authorization: Bearer ${COHERE_API_KEY}" 
  -H "Content-Type: application/json" 
  -H "Accept: application/json" 
  -d '{
    "model": "command-a-plus-05-2026",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain retrieval-augmented generation in two sentences."}
    ]
  }'

The same request can be made with Cohere's current Python SDK. Install the package with pip install -U cohere and use the v2 client.

import os
import cohere

client = cohere.ClientV2(api_key=os.environ["COHERE_API_KEY"])

response = client.chat(
    model="command-a-plus-05-2026",
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "Explain semantic search in two sentences."}
    ],
)

text = "".join(
    part.text for part in response.message.content
    if part.type == "text"
)
print(text)

How should you read the response?

A Chat API v2 response contains a message with content parts. Text content is represented as a content part whose type is text. Applications should collect the text parts rather than assuming the response is always one plain string, because a response can also contain tool calls, citations or other event types depending on the request.

For production code, handle unsuccessful HTTP responses, empty or unexpected content, rate-limit responses and model-specific errors. Preserve usage metadata when available so that your application can measure token consumption and investigate cost or latency changes.

How does Cohere API pricing work?

Cohere uses usage-based pricing that varies by model and endpoint. Generative models are generally priced by input and output tokens. Reranking is priced by search volume or applicable units, while embeddings are priced by embedded tokens or workload units. Dedicated or managed deployment options have separate pricing.

Trial keys are free but rate-limited. The supplied rate-limit information describes trial keys as generally limited to 1,000 API calls per month and commonly 20 requests per minute for Chat models. Production limits vary by model and endpoint; examples in the supplied documentation include as many as 500 requests per minute for several active Chat models, 2,000 inputs per minute for Embed and 1,000 requests per minute for Rerank. Some newer models require contacting sales for production limits.

Check Cohere's current pricing and rate-limit documentation before estimating costs. Do not assume that a limit or price for one model applies to another.

What can developers build with the API?

Streaming responses

Chat responses can be streamed as server-sent events. Instead of waiting for the complete answer, the application receives events as content is generated. This is useful for user interfaces where displaying the first words quickly matters. Stream events may include message starts, content deltas, tool calls, citations and completion events.

stream = client.chat_stream(
    model="command-a-plus-05-2026",
    messages=[
        {"role": "user", "content": "Explain embeddings in three short steps."}
    ],
)

for event in stream:
    if event.type == "content-delta" and event.delta and event.delta.message:
        print(event.delta.message, end="", flush=True)
print()

Embeddings, reranking and RAG

A typical RAG pipeline uses Embed to represent documents and user queries as vectors, a search system to find candidate passages, and Rerank to improve the ordering of those candidates. The selected passages are then supplied to Chat as context. This separates retrieval from generation and lets the application control which enterprise information is shown to the model.

Cohere also supports citations and retrieval documents in Chat workflows. Your application should still enforce access permissions before sending retrieved content to the model.

Tool and function calling

Tools let the model request an action such as querying an internal system or looking up an order. The model does not execute the tool itself. Your application receives a tool call, validates the requested arguments, performs the action, and sends the result back in a follow-up message.

Tool definitions use a JSON Schema-like description. Cohere supports multi-step and parallel tool patterns, but the application remains responsible for authorization, validation, retries, timeouts and preventing unsafe actions. The experimental strict_tools option can enforce tool names, required parameters and parameter types on compatible Chat API v2 workflows.

Structured outputs

Chat API v2 supports the response_format parameter for JSON mode and JSON Schema-style output on compatible models. For example:

response = client.chat(
    model="command-a-plus-05-2026",
    messages=[
        {"role": "user", "content": "Describe semantic search with fields name and benefit."}
    ],
    response_format={"type": "json_object"},
)

print(response.message.content[0].text)

JSON mode does not remove the need to validate the result. Model compatibility varies, and structured output should be treated as a response-format feature rather than a separate JSON-only API.

Images, documents and files

Compatible vision models accept images through HTTP URLs or base64 data URLs. Supported image formats include PNG, JPEG, WEBP and non-animated GIF, subject to documented image-count and request-size limits.

Cohere also supports CSV and JSONL dataset uploads for embedding jobs and provides document and image parsing through the Parse API. These uploads are task-specific; they are not a universal, persistent file-search store. Dataset files are automatically deleted after the documented retention period of 30 days unless deleted sooner.

Which SDKs and developer tools are available?

Cohere officially supports Python, TypeScript, Java and Go SDKs. REST and HTTP access can be used from other languages, including PHP. The current Python client uses cohere.ClientV2; the current TypeScript package uses CohereClientV2 from cohere-ai.

Use the official API reference, model documentation, SDK repositories, playground, cookbooks and deprecation notices together. The playground is useful for exploration, while automated tests should verify the exact request and response structures your application depends on.

What limits and production issues matter?

  • Rate limits: Limits differ between trial and production keys and between models and endpoints. Implement backoff and handle rate-limit responses.
  • Latency: There is no single universal latency guarantee. Model choice, prompt size, output length, image detail, tools, queue priority and deployment affect response time. Streaming can reduce perceived waiting time.
  • Model compatibility: Tools, structured output and image input are not necessarily available on every model. Test each capability with the selected model.
  • Data retention: Retention depends on the product and deployment. Model Vault can support Zero Data Retention configurations, while private and partner-cloud deployments can provide stronger customer-controlled data boundaries.
  • Security: Keep API keys server-side, validate tool arguments, filter retrieved content and apply application-level authorization before exposing enterprise data or actions.
  • Deprecations: Legacy endpoints such as /v1/generate, /v1/summarize and /v1/classify, along with related legacy features, should not be the basis of new applications.
  • Fine-tuning: Cohere's deprecation documentation states that the listed platform fine-tuning capabilities were retired effective September 15, 2025. Do not plan a new integration around historical fine-tuning endpoints.

Is Cohere API a good choice for your project?

Cohere is a strong candidate when an organization needs multilingual generation, semantic retrieval, reranking, document processing, agent workflows or deployment options that include private, dedicated, cloud or air-gapped environments. The separation between Chat, Embed, Rerank and Parse is useful for building controlled enterprise search and RAG systems.

It may be a poor choice if the main requirement is a free consumer chatbot, first-party mobile or desktop applications, image or video generation, or a fully managed personal productivity ecosystem. Enterprise pricing is often custom, advanced workflows require application-side engineering, and capabilities can vary by model and deployment. Evaluate the complete workflow—including data governance, retrieval quality, operational limits and tool safety—rather than testing text generation alone.

Sources 21