Command A

Command A Reasoning

by Cohere · Live

Cohere Command A Reasoning is a 111-billion-parameter enterprise reasoning model released in August 2025. It combines configurable thinking, a 256K-token context window, 32K-token output, tool use, structured outputs, multilingual support across 23 languages, and hosted or Model Vault deployment options. It is best suited to complex agents, RAG, long-context analysis, and workflow automation rather than multimodal input or low-latency routine generation.

Text Reasoning Coding
Command A Reasoning is Cohere’s first dedicated reasoning model. It combines configurable internal thinking with a 256,000-token context window, 32,000-token maximum output, tool calling, structured outputs, and support for 23 languages. Its main role is handling complex, multi-step enterprise tasks rather than serving as a low-latency general chatbot or a media-generation model.
Outputs

What Command A Reasoning can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Structured output
Model profile

Performance characteristics

8/10 Reasoning
7/10 Coding
5/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family Command A
Model type Reasoning
Context window 256K tokens
Maximum output 32K tokens
Knowledge cutoff 2024-06-01
Release date 2025-08-21
Status Live
Knowledge cutoff notes

Cohere documents June 1, 2024 as the underlying knowledge cutoff. Web search, retrieval, connectors, and developer-provided tools can supply newer external information during use but do not change the model’s cutoff.

Model notes

Cohere’s first dedicated reasoning model, released August 21, 2025. The exact model ID is command-a-reasoning-08-2025. It has 111 billion parameters, supports configurable thinking and user-controlled thinking budgets, and is optimized for four H100 GPUs in production or four A100 GPUs for evaluation. Cohere documents support for 23 languages, tool use, citations, structured outputs, and Chat API access. Cohere lists the hosted API as free until rate limits are reached, while newer model variants may require contacting Cohere for production access. Model Vault pricing is per instance-hour: $48/hour for L and $57.50/hour for XL. Reasoning, coding, speed, and cost scores are editorial comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input Free until applicable API rate limits; private Model Vault deployment is billed per instance-hour rather than per token. Standard Vault pricing lists $48/hour for L and $57.50/hour for XL performance tiers.
Output Free until applicable API rate limits on Cohere-hosted trial and production access; Model Vault pricing is instance-hour based rather than output-token based.
Model guide

Command A Reasoning: Cohere’s Long-Context Model for Enterprise Agents

Cohere Command A Reasoning is a 111-billion-parameter hybrid reasoning model released on August 21, 2025. It is designed for complex enterprise agents, retrieval-augmented generation, multilingual analysis, tool use, and long-context workflows, with a 256,000-token context window, up to 32,000 output tokens, configurable thinking, and support for 23 languages.

What is Command A Reasoning?

Command A Reasoning is Cohere’s first dedicated reasoning model and part of the Command A model family. Its canonical model identifier is command-a-reasoning-08-2025. Cohere released it on August 21, 2025, and the supplied documentation identifies it as a live model.

The model is intended for applications that must work through several steps before responding or taking an action. Typical examples include enterprise agents that search internal information, compare documents, call business systems, retrieve supporting evidence, and then produce a grounded answer or complete a workflow.

Unlike a conventional text-generation model that generally moves directly from a prompt to an answer, Command A Reasoning can use a configurable thinking operation. When enabled, the model can produce a thinking block before its final text response. Developers can control whether thinking is enabled and adjust the available thinking budget, allowing them to trade additional reasoning effort against response speed and resource use.

Where it fits in Cohere’s lineup

Command A Reasoning sits in Cohere’s Command family, which focuses on text generation, agents, and enterprise applications. It complements rather than replaces Cohere’s other model families, including Embed and Rerank for retrieval systems, Parse for document intelligence, and Aya for multilingual and multimodal research use cases.

The important distinction is its specialization. Command A Reasoning is aimed at difficult reasoning and agentic tasks, while a simpler generation model may be preferable when an application mainly needs fast, inexpensive text responses. The model is also separate from Cohere’s enterprise product layer: it can be accessed through Cohere’s Chat APIs and can be deployed privately through Model Vault, while applications may use it inside broader retrieval or workflow systems.

Verified technical specifications

SpecificationCommand A Reasoning
ProviderCohere
Model IDcommand-a-reasoning-08-2025
Model familyCommand A
Model typeHybrid reasoning model
Parameters111 billion
Release dateAugust 21, 2025
Context window256,000 tokens
Maximum output32,000 tokens
Knowledge cutoffJune 1, 2024
InputText
OutputText, with optional thinking content
Languages23 languages, according to Cohere
API accessCohere Chat API, including Chat API v2, Chat Completions, and Chat API v1 interfaces

The 256,000-token context window is useful for large document collections, lengthy case files, multi-step conversations, and retrieval results that would exceed the limits of many smaller models. The 32,000-token output limit also allows long reports or detailed structured responses, although producing very long answers is not always the most efficient design for a production application.

The June 1, 2024 knowledge cutoff is separate from live retrieval. Search results, connected data, or developer-provided tools can supply newer information during a request, but they do not update the model’s underlying training knowledge.

Reasoning, tool use, and retrieval

Command A Reasoning is designed to support multi-step workflows rather than only single-turn question answering. A developer can define tools representing actions such as enterprise search, database queries, API requests, or business operations. The model can then request a tool call, receive the returned information, and continue reasoning before producing a response or requesting another action.

This makes the model suitable for retrieval-augmented generation, commonly abbreviated as RAG. In a RAG system, an application retrieves relevant passages from a document store or business database and supplies them to the model. Command A Reasoning can use those passages as evidence while composing an answer, report, or recommendation.

Tool use does not mean the model independently has unrestricted access to the web or a company’s systems. The application still defines the available tools, executes them, controls permissions, and returns their results. Cohere’s documentation supports tool use and citations, while web search should be treated as an external platform capability rather than as a change to the model’s knowledge cutoff.

Structured output and developer controls

The model supports structured machine-readable output through Cohere’s Chat API. This is useful when the response must be consumed by software rather than read only by a person. For example, an application could request a consistent object containing a case summary, risk categories, cited sources, and recommended next actions.

Command A Reasoning also supports streaming, allowing an application to receive response content progressively instead of waiting for the entire result. Thinking can be enabled or disabled, and the available thinking budget can be controlled. These settings give developers a way to balance reasoning depth, latency, and operational cost for different requests.

The supplied research does not verify a distinct JSON-mode flag for this exact model. It does verify structured outputs, so structured-output support should not automatically be described as a separate JSON mode.

Input and output modalities

Command A Reasoning is a text-in, text-out model. It accepts text and returns text, including optional reasoning content when thinking is enabled. The supplied specifications do not identify native image, audio, or video input, and the model does not generate images, audio, video, music, or speech.

This limitation matters when selecting a model for multimodal work. A document-analysis pipeline can still use Command A Reasoning after another service extracts text from a PDF or transcribes audio, but the model itself should not be presented as directly understanding those media types. Applications that need native image understanding, audio processing, or media generation require a more appropriate specialized model or pipeline.

Coding and performance trade-offs

Command A Reasoning can assist with coding as part of its general text generation and tool-use capabilities. It may be useful for explaining code, planning implementation steps, reviewing technical material, generating scripts, or coordinating tools that interact with development systems. The supplied research does not provide an official coding benchmark or a provider-published coding score.

Its main performance advantage is depth and context capacity, not necessarily minimum latency. Reasoning steps, large prompts, tool calls, and long outputs can all increase the time and resources required for a response. The editorial assessment supplied with this model rates reasoning at 8 out of 10, coding at 7 out of 10, speed at 5 out of 10, and cost at 4 out of 10. These are comparative editorial estimates, not Cohere benchmarks or guaranteed service levels.

Cohere states that the model is optimized for production deployment on four H100 GPUs. Four A100 GPUs are described as suitable for non-production evaluation workloads. Those infrastructure requirements help explain why the model is better aligned with enterprise deployments than with small, inexpensive local applications.

Pricing and availability

Cohere’s hosted API documentation describes Command A Reasoning as free for trial and production keys until applicable rate limits are reached. This does not mean unlimited free production use: access remains subject to API limits, and newer model variants may require contacting Cohere for production rate limits or approval.

For private deployment, Cohere lists the model in Model Vault. Standard Vault pricing is instance-based rather than token-based. The documented rates are $48 per hour for the L performance tier and $57.50 per hour for the XL performance tier. Generative Model Vault access may require a waitlist or commercial approval.

These two pricing paths should not be compared as if they were equivalent per-token plans. Hosted API access is governed by keys and rate limits, while Model Vault charges for provisioned performance capacity. Organizations considering private deployment should account for utilization, infrastructure requirements, security needs, and commercial terms rather than estimating cost only from the number of generated tokens.

Best use cases

  • Complex enterprise agents: Multi-step assistants that need to reason, call tools, inspect results, and continue an action sequence.
  • Retrieval-augmented generation: Answers and reports grounded in large internal document collections or business data.
  • Long-context analysis: Reviewing extensive contracts, research collections, case files, or technical documentation within one request.
  • Multilingual operations: Research, support, analysis, or workflow automation across the 23 languages Cohere identifies as supported.
  • Structured business workflows: Producing predictable machine-readable results for downstream software.
  • Private enterprise deployments: Use cases where Model Vault or other controlled deployment options are more appropriate than a standard public API.

When to choose Command A Reasoning

Choose Command A Reasoning when the application benefits from deliberate multi-step reasoning, a very large context window, tool interaction, multilingual support, or enterprise deployment controls. It is especially relevant when a response must combine retrieved evidence with several decisions or actions instead of simply generating a short answer.

A faster, smaller generation model may be a better choice for high-volume classification, short customer replies, simple extraction, or latency-sensitive interactions. A retrieval-specific model may be more suitable when the main problem is finding or ranking documents rather than reasoning over them. A model with native multimodal input is preferable when users need to submit images, audio, or video directly. Finally, organizations that need predictable token-based pricing should compare Cohere’s hosted terms carefully with alternatives, because Model Vault uses instance-hour pricing.

Command A Reasoning is therefore best understood as a deliberate reasoning engine for enterprise workflows, not as an all-purpose media model or the lowest-cost option for routine text generation. Its value comes from combining configurable thinking, long context, tool use, structured responses, and Cohere’s private deployment options in one model.


Answers to Frequently Asked Questions

How much does Command A Reasoning cost?
Cohere’s hosted API documentation describes the model as free for trial and production keys until applicable rate limits are reached, so usage is not unlimited. Private Model Vault deployment is instance-based, with documented rates of $48 per hour for the L tier and $57.50 per hour for the XL tier; commercial approval or a waitlist may apply.
Does Command A Reasoning support images, audio, or video?
No. Command A Reasoning is a text-in, text-out model with optional thinking content. It does not provide native image, audio, or video understanding or generation, although other services can preprocess those media types into text for the model.
Can Command A Reasoning use tools and retrieval-augmented generation?
Yes. Developers can connect tools for enterprise search, database queries, API calls, and business operations. The model can use retrieved passages or tool results as evidence during multi-step reasoning, while the application controls tool execution, permissions, and access to external systems.
What is Command A Reasoning?
Command A Reasoning is Cohere’s dedicated hybrid reasoning model for multi-step enterprise tasks. It can analyze information, use developer-defined tools, retrieve evidence, and produce grounded answers or workflow actions. Its model ID is command-a-reasoning-08-2025.
What are the context window and output limits of Command A Reasoning?
Command A Reasoning supports a 256,000-token context window and a maximum output of 32,000 tokens. This makes it suitable for long documents, extensive case files, large retrieval results, and detailed reports.


Sources 10
Provider

About Cohere