What is Command A Reasoning?
Command A Reasoning is Cohere’s first dedicated reasoning model and part of the Command A model family. Its canonical model identifier is command-a-reasoning-08-2025. Cohere released it on August 21, 2025, and the supplied documentation identifies it as a live model.
The model is intended for applications that must work through several steps before responding or taking an action. Typical examples include enterprise agents that search internal information, compare documents, call business systems, retrieve supporting evidence, and then produce a grounded answer or complete a workflow.
Unlike a conventional text-generation model that generally moves directly from a prompt to an answer, Command A Reasoning can use a configurable thinking operation. When enabled, the model can produce a thinking block before its final text response. Developers can control whether thinking is enabled and adjust the available thinking budget, allowing them to trade additional reasoning effort against response speed and resource use.
Where it fits in Cohere’s lineup
Command A Reasoning sits in Cohere’s Command family, which focuses on text generation, agents, and enterprise applications. It complements rather than replaces Cohere’s other model families, including Embed and Rerank for retrieval systems, Parse for document intelligence, and Aya for multilingual and multimodal research use cases.
The important distinction is its specialization. Command A Reasoning is aimed at difficult reasoning and agentic tasks, while a simpler generation model may be preferable when an application mainly needs fast, inexpensive text responses. The model is also separate from Cohere’s enterprise product layer: it can be accessed through Cohere’s Chat APIs and can be deployed privately through Model Vault, while applications may use it inside broader retrieval or workflow systems.
Verified technical specifications
| Specification | Command A Reasoning |
|---|---|
| Provider | Cohere |
| Model ID | command-a-reasoning-08-2025 |
| Model family | Command A |
| Model type | Hybrid reasoning model |
| Parameters | 111 billion |
| Release date | August 21, 2025 |
| Context window | 256,000 tokens |
| Maximum output | 32,000 tokens |
| Knowledge cutoff | June 1, 2024 |
| Input | Text |
| Output | Text, with optional thinking content |
| Languages | 23 languages, according to Cohere |
| API access | Cohere Chat API, including Chat API v2, Chat Completions, and Chat API v1 interfaces |
The 256,000-token context window is useful for large document collections, lengthy case files, multi-step conversations, and retrieval results that would exceed the limits of many smaller models. The 32,000-token output limit also allows long reports or detailed structured responses, although producing very long answers is not always the most efficient design for a production application.
The June 1, 2024 knowledge cutoff is separate from live retrieval. Search results, connected data, or developer-provided tools can supply newer information during a request, but they do not update the model’s underlying training knowledge.
Reasoning, tool use, and retrieval
Command A Reasoning is designed to support multi-step workflows rather than only single-turn question answering. A developer can define tools representing actions such as enterprise search, database queries, API requests, or business operations. The model can then request a tool call, receive the returned information, and continue reasoning before producing a response or requesting another action.
This makes the model suitable for retrieval-augmented generation, commonly abbreviated as RAG. In a RAG system, an application retrieves relevant passages from a document store or business database and supplies them to the model. Command A Reasoning can use those passages as evidence while composing an answer, report, or recommendation.
Tool use does not mean the model independently has unrestricted access to the web or a company’s systems. The application still defines the available tools, executes them, controls permissions, and returns their results. Cohere’s documentation supports tool use and citations, while web search should be treated as an external platform capability rather than as a change to the model’s knowledge cutoff.
Structured output and developer controls
The model supports structured machine-readable output through Cohere’s Chat API. This is useful when the response must be consumed by software rather than read only by a person. For example, an application could request a consistent object containing a case summary, risk categories, cited sources, and recommended next actions.
Command A Reasoning also supports streaming, allowing an application to receive response content progressively instead of waiting for the entire result. Thinking can be enabled or disabled, and the available thinking budget can be controlled. These settings give developers a way to balance reasoning depth, latency, and operational cost for different requests.
The supplied research does not verify a distinct JSON-mode flag for this exact model. It does verify structured outputs, so structured-output support should not automatically be described as a separate JSON mode.
Input and output modalities
Command A Reasoning is a text-in, text-out model. It accepts text and returns text, including optional reasoning content when thinking is enabled. The supplied specifications do not identify native image, audio, or video input, and the model does not generate images, audio, video, music, or speech.
This limitation matters when selecting a model for multimodal work. A document-analysis pipeline can still use Command A Reasoning after another service extracts text from a PDF or transcribes audio, but the model itself should not be presented as directly understanding those media types. Applications that need native image understanding, audio processing, or media generation require a more appropriate specialized model or pipeline.
Coding and performance trade-offs
Command A Reasoning can assist with coding as part of its general text generation and tool-use capabilities. It may be useful for explaining code, planning implementation steps, reviewing technical material, generating scripts, or coordinating tools that interact with development systems. The supplied research does not provide an official coding benchmark or a provider-published coding score.
Its main performance advantage is depth and context capacity, not necessarily minimum latency. Reasoning steps, large prompts, tool calls, and long outputs can all increase the time and resources required for a response. The editorial assessment supplied with this model rates reasoning at 8 out of 10, coding at 7 out of 10, speed at 5 out of 10, and cost at 4 out of 10. These are comparative editorial estimates, not Cohere benchmarks or guaranteed service levels.
Cohere states that the model is optimized for production deployment on four H100 GPUs. Four A100 GPUs are described as suitable for non-production evaluation workloads. Those infrastructure requirements help explain why the model is better aligned with enterprise deployments than with small, inexpensive local applications.
Pricing and availability
Cohere’s hosted API documentation describes Command A Reasoning as free for trial and production keys until applicable rate limits are reached. This does not mean unlimited free production use: access remains subject to API limits, and newer model variants may require contacting Cohere for production rate limits or approval.
For private deployment, Cohere lists the model in Model Vault. Standard Vault pricing is instance-based rather than token-based. The documented rates are $48 per hour for the L performance tier and $57.50 per hour for the XL performance tier. Generative Model Vault access may require a waitlist or commercial approval.
These two pricing paths should not be compared as if they were equivalent per-token plans. Hosted API access is governed by keys and rate limits, while Model Vault charges for provisioned performance capacity. Organizations considering private deployment should account for utilization, infrastructure requirements, security needs, and commercial terms rather than estimating cost only from the number of generated tokens.
Best use cases
- Complex enterprise agents: Multi-step assistants that need to reason, call tools, inspect results, and continue an action sequence.
- Retrieval-augmented generation: Answers and reports grounded in large internal document collections or business data.
- Long-context analysis: Reviewing extensive contracts, research collections, case files, or technical documentation within one request.
- Multilingual operations: Research, support, analysis, or workflow automation across the 23 languages Cohere identifies as supported.
- Structured business workflows: Producing predictable machine-readable results for downstream software.
- Private enterprise deployments: Use cases where Model Vault or other controlled deployment options are more appropriate than a standard public API.
When to choose Command A Reasoning
Choose Command A Reasoning when the application benefits from deliberate multi-step reasoning, a very large context window, tool interaction, multilingual support, or enterprise deployment controls. It is especially relevant when a response must combine retrieved evidence with several decisions or actions instead of simply generating a short answer.
A faster, smaller generation model may be a better choice for high-volume classification, short customer replies, simple extraction, or latency-sensitive interactions. A retrieval-specific model may be more suitable when the main problem is finding or ranking documents rather than reasoning over them. A model with native multimodal input is preferable when users need to submit images, audio, or video directly. Finally, organizations that need predictable token-based pricing should compare Cohere’s hosted terms carefully with alternatives, because Model Vault uses instance-hour pricing.
Command A Reasoning is therefore best understood as a deliberate reasoning engine for enterprise workflows, not as an all-purpose media model or the lowest-cost option for routine text generation. Its value comes from combining configurable thinking, long context, tool use, structured responses, and Cohere’s private deployment options in one model.

