What is Cohere Command A?
Cohere Command A is a 111-billion-parameter large language model from Cohere, released on March 13, 2025. It is currently listed as a live model in Cohere's catalog and is available through the Cohere API and selected deployment platforms.
The model is built primarily for enterprise language workloads rather than consumer chat. Its intended uses include retrieval-augmented generation (RAG), long-context document analysis, multilingual applications, tool-connected assistants, structured text generation and agentic workflows. In practical terms, Command A can read information supplied in a prompt or retrieved from an organization's systems, decide when an external tool is needed, and return a response in a format suitable for downstream software.
Command A is a text model. The supplied specifications identify text input and text output, with no native image, audio or video input or output. External systems can still give an application access to other modalities or services, but those integrations should not be confused with native multimodal support from Command A itself.
Command A specifications at a glance
| Specification | Details |
|---|---|
| Provider | Cohere |
| Release date | March 13, 2025 |
| Status | Live |
| Model size | 111 billion parameters |
| Context window | 256,000 tokens |
| Maximum output | 8,000 tokens |
| Knowledge cutoff | June 1, 2024 |
| Languages | 23 supported languages |
| Input and output | Text input and text output |
| Tool use | Supported |
| Streaming | Supported |
| API input price | $2.50 per 1 million tokens |
| API output price | $10.00 per 1 million tokens |
These are the supplied provider and catalog specifications. Comparative ratings sometimes associated with the model should be treated separately: the research gives Command A editorial scores of 8 out of 10 for reasoning, coding and speed, and 6 out of 10 for cost. Those ratings are subjective comparisons, not scores published by Cohere.
Why the 256K context window matters
Command A has a 256,000-token context window, which is the amount of text the model can consider across the prompt and conversation context. This is substantially more useful for enterprise document work than a short-context model because an application can supply lengthy reports, collections of records, policy material or retrieved passages in one request.
The context window does not mean that the model automatically knows every company document or can search the internet. A RAG application must still retrieve relevant information and include it in the model request. Command A can then use that supplied material to draft an answer, summarize evidence, compare documents or produce a response with citations when the surrounding tool workflow provides them.
The model's knowledge cutoff is June 1, 2024. Information created after that date is not part of the underlying training knowledge described in the supplied documentation. Retrieval, an external database or a connected search tool can provide newer information during a request, but that does not change the model's knowledge cutoff.
Tool use, citations and agentic workflows
Command A supports tool use, allowing an application to describe functions or services that the model can call as part of a task. A tool might retrieve records, query an internal system, perform a calculation or trigger another approved operation. The model itself does not become a database or an autonomous business system; the surrounding application remains responsible for implementing tools, enforcing permissions and deciding whether an action should be executed.
This makes Command A suitable for agentic workflows in which the model performs several language and tool steps instead of returning only one standalone answer. For example, an enterprise assistant could receive a question, request relevant documents from a retrieval service, inspect the returned material and draft a structured response. Cohere's documentation also covers citations for tool use and streaming for tool-use interactions.
Tool use does not establish native web search. Command A is not documented here as having built-in first-party web browsing or current real-time information. An application can connect it to an external search engine, but the search capability comes from that external tool and its implementation.
Reasoning, coding and structured output
Command A is positioned for multi-step enterprise tasks, including document analysis, retrieval workflows and agent orchestration. The supplied editorial assessment rates its reasoning capability at 8 out of 10 and its coding capability at 8 out of 10. These ratings indicate a strong comparative fit for reasoning-heavy and code-related workloads, but they are editorial estimates rather than provider benchmarks.
Its coding usefulness is particularly relevant when code generation is part of a larger workflow: generating queries, transforming data, creating automation snippets or helping an agent interact with business systems. The research does not provide benchmark results or language-specific coding scores, so performance should be validated against the application's own examples.
Command A supports structured outputs. This allows an application to request text that follows a defined structure rather than relying entirely on free-form prose. Structured responses are useful for extracting fields from documents, returning classifications, producing workflow parameters or passing model results to software. Schema validation and error handling are still application responsibilities; structured output support does not guarantee that every response will be correct or that the requested schema fits every task.
Multilingual and enterprise-oriented use
Cohere documents support for 23 languages, making Command A relevant to organizations that need one model for multilingual search, analysis, generation or customer and employee workflows. The exact quality of a multilingual result can vary by language and task, so teams should test important production languages rather than assuming equal performance across all 23.
The model's overall positioning is enterprise-focused. Command A is a strong candidate for systems that need long documents, retrieval, tool calls, citations and controlled text formats in the same workflow. It can fit applications such as financial text processing, policy and contract analysis, internal knowledge assistants, research automation, multilingual support operations and document-based agents.
Command A pricing and cost trade-offs
The supplied API pricing is $2.50 per 1 million input tokens and $10 per 1 million output tokens. Input tokens are the material sent to the model, while output tokens are the text it generates. Because output is priced at four times the input rate, applications can reduce costs by limiting unnecessary response length, selecting concise instructions and avoiding repeated generation when a shorter result is sufficient.
Command A is not positioned as the lowest-cost option for very high-volume, routine text generation. The editorial cost score is 6 out of 10, reflecting a trade-off between its broad enterprise capabilities and its per-token price. Its value is more apparent when a task benefits from a large context window, tool use, multilingual handling or a reliable structured workflow. A smaller or cheaper model may be more appropriate for simple classification, short summaries or large batches of low-complexity text.
The model has an 8,000-token maximum output. That is enough for substantial reports, extracted records and detailed answers, but it is not unlimited. Applications generating long documents should handle truncation, divide work into stages or use an application-level process that combines multiple responses.
Main strengths and limitations
Strengths
- Very large context: The 256K-token window supports lengthy source material and broad retrieval contexts.
- Enterprise workflow fit: Tool use, citations, structured outputs and streaming support agentic applications.
- Multilingual coverage: Support for 23 languages is useful for organizations serving multiple regions.
- Strong text reasoning profile: The editorial assessment rates reasoning and coding at 8 out of 10, although these are not provider-published benchmark scores.
- Production-oriented availability: The model is listed as live and is available through Cohere's API and selected deployment platforms.
Limitations
- Text only: Command A does not natively accept or generate images, audio or video.
- No built-in web search: Current information requires retrieval or an external search tool.
- Knowledge cutoff: The underlying model cutoff is June 1, 2024.
- Cost: Output pricing is higher than input pricing, and the model may not be economical for simple high-volume generation.
- Output ceiling: Responses are limited to 8,000 tokens, so very large documents may require multiple stages.
- Integration requirements: Tool use, retrieval, permissions, citations and validation must be implemented and managed by the surrounding application.
When to choose Command A
Choose Command A when the application needs to combine long-context text processing with enterprise retrieval, tools or structured responses. It is especially suitable when a single workflow must inspect large documents, use external business data, support multiple languages and return information that software can consume.
It is also a reasonable choice for teams building document-focused agents rather than a simple chat interface. The 256K context window can reduce the need to split source material into many small requests, while tool use allows the model to work with systems outside the model itself.
Another model may be more appropriate when the priority is the lowest possible cost, extremely high-volume short-form generation, native image or audio processing, or built-in current web search. A smaller model can also be preferable when the task does not need Command A's long context, multilingual scope or agent features. For current facts, Command A should be paired with a suitable retrieval or search system rather than used alone.
Overall assessment
Cohere Command A is a specialized enterprise language model whose main advantage is the combination of a 256K context window, tool use, structured outputs, multilingual support and agent-oriented design. It is best understood as a foundation for document and workflow applications, not as a consumer multimodal assistant or a standalone source of current information.
Its pricing and text-only design limit its suitability for some workloads, but those trade-offs can be justified when the application needs to process substantial context and connect model reasoning to retrieval or business tools. Teams evaluating it should test their own documents, languages, tool sequences and output schemas, while treating the published limits and pricing as the verified baseline.

