Command A

Command A

by Cohere · Live

Cohere Command A is a live 111-billion-parameter enterprise language model for long-context document analysis, RAG, multilingual applications, tool use and agentic workflows. It supports a 256K-token context window, 8K-token maximum output, structured outputs, streaming and 23 languages. The model is text-only, has a June 1, 2024 knowledge cutoff, and costs $2.50 per million input tokens and $10 per million output tokens.

Text Reasoning Coding
Cohere Command A is a text-only language model aimed at enterprise applications that need to process large documents, retrieve information, call external tools, and produce structured responses. It combines a 256K-token context window with tool use, citations, structured outputs, multilingual support and an 8K-token maximum response. The model is available through Cohere's API and selected deployment platforms, with pricing that favors focused, high-value workloads over the cheapest possible high-volume generation.
Outputs

What Command A can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming JSON mode Structured output
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
8/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Command A
Model type General Purpose
Context window 256K tokens
Maximum output 8K tokens
Knowledge cutoff June 1, 2024
Release date March 13, 2025
Status Live
Knowledge cutoff notes

Cohere's exact Command A model documentation explicitly lists the knowledge cutoff as June 1, 2024. External tools, retrieval, citations or search integrations can provide newer information during use but do not change the underlying model cutoff.

Model notes

Command A is Cohere's 111-billion-parameter enterprise language model. Cohere documents a 256K context window, 8K maximum output, June 1, 2024 knowledge cutoff, support for 23 languages, tool use, agents, RAG, citations and structured outputs. It is text-only for model input and output. The model can be connected to external search engines or other tools through tool use, but this does not establish a native first-party web-search capability. It is listed as Live in Cohere's current model catalog. Cohere documents deployment on two A100 or H100 GPUs and reports 150% higher throughput than Command R+ 08-2024. Editorial scores are comparative estimates rather than vendor-provided ratings.

Cost

Model pricing

Input $2.50 per 1 million tokens
Output $10.00 per 1 million tokens
Model guide

Cohere Command A: Long-Context Enterprise AI for RAG and Agents

Cohere Command A is a 111-billion-parameter enterprise language model released on March 13, 2025. It is designed for long-context generation, multilingual applications, retrieval-augmented generation, tool use, structured text generation, and agentic workflows. Its defining specifications are a 256,000-token context window, an 8,000-token maximum output, support for 23 languages, and API pricing of $2.50 per million input tokens and $10 per million output tokens.

What is Cohere Command A?

Cohere Command A is a 111-billion-parameter large language model from Cohere, released on March 13, 2025. It is currently listed as a live model in Cohere's catalog and is available through the Cohere API and selected deployment platforms.

The model is built primarily for enterprise language workloads rather than consumer chat. Its intended uses include retrieval-augmented generation (RAG), long-context document analysis, multilingual applications, tool-connected assistants, structured text generation and agentic workflows. In practical terms, Command A can read information supplied in a prompt or retrieved from an organization's systems, decide when an external tool is needed, and return a response in a format suitable for downstream software.

Command A is a text model. The supplied specifications identify text input and text output, with no native image, audio or video input or output. External systems can still give an application access to other modalities or services, but those integrations should not be confused with native multimodal support from Command A itself.

Command A specifications at a glance

SpecificationDetails
ProviderCohere
Release dateMarch 13, 2025
StatusLive
Model size111 billion parameters
Context window256,000 tokens
Maximum output8,000 tokens
Knowledge cutoffJune 1, 2024
Languages23 supported languages
Input and outputText input and text output
Tool useSupported
StreamingSupported
API input price$2.50 per 1 million tokens
API output price$10.00 per 1 million tokens

These are the supplied provider and catalog specifications. Comparative ratings sometimes associated with the model should be treated separately: the research gives Command A editorial scores of 8 out of 10 for reasoning, coding and speed, and 6 out of 10 for cost. Those ratings are subjective comparisons, not scores published by Cohere.

Why the 256K context window matters

Command A has a 256,000-token context window, which is the amount of text the model can consider across the prompt and conversation context. This is substantially more useful for enterprise document work than a short-context model because an application can supply lengthy reports, collections of records, policy material or retrieved passages in one request.

The context window does not mean that the model automatically knows every company document or can search the internet. A RAG application must still retrieve relevant information and include it in the model request. Command A can then use that supplied material to draft an answer, summarize evidence, compare documents or produce a response with citations when the surrounding tool workflow provides them.

The model's knowledge cutoff is June 1, 2024. Information created after that date is not part of the underlying training knowledge described in the supplied documentation. Retrieval, an external database or a connected search tool can provide newer information during a request, but that does not change the model's knowledge cutoff.

Tool use, citations and agentic workflows

Command A supports tool use, allowing an application to describe functions or services that the model can call as part of a task. A tool might retrieve records, query an internal system, perform a calculation or trigger another approved operation. The model itself does not become a database or an autonomous business system; the surrounding application remains responsible for implementing tools, enforcing permissions and deciding whether an action should be executed.

This makes Command A suitable for agentic workflows in which the model performs several language and tool steps instead of returning only one standalone answer. For example, an enterprise assistant could receive a question, request relevant documents from a retrieval service, inspect the returned material and draft a structured response. Cohere's documentation also covers citations for tool use and streaming for tool-use interactions.

Tool use does not establish native web search. Command A is not documented here as having built-in first-party web browsing or current real-time information. An application can connect it to an external search engine, but the search capability comes from that external tool and its implementation.

Reasoning, coding and structured output

Command A is positioned for multi-step enterprise tasks, including document analysis, retrieval workflows and agent orchestration. The supplied editorial assessment rates its reasoning capability at 8 out of 10 and its coding capability at 8 out of 10. These ratings indicate a strong comparative fit for reasoning-heavy and code-related workloads, but they are editorial estimates rather than provider benchmarks.

Its coding usefulness is particularly relevant when code generation is part of a larger workflow: generating queries, transforming data, creating automation snippets or helping an agent interact with business systems. The research does not provide benchmark results or language-specific coding scores, so performance should be validated against the application's own examples.

Command A supports structured outputs. This allows an application to request text that follows a defined structure rather than relying entirely on free-form prose. Structured responses are useful for extracting fields from documents, returning classifications, producing workflow parameters or passing model results to software. Schema validation and error handling are still application responsibilities; structured output support does not guarantee that every response will be correct or that the requested schema fits every task.

Multilingual and enterprise-oriented use

Cohere documents support for 23 languages, making Command A relevant to organizations that need one model for multilingual search, analysis, generation or customer and employee workflows. The exact quality of a multilingual result can vary by language and task, so teams should test important production languages rather than assuming equal performance across all 23.

The model's overall positioning is enterprise-focused. Command A is a strong candidate for systems that need long documents, retrieval, tool calls, citations and controlled text formats in the same workflow. It can fit applications such as financial text processing, policy and contract analysis, internal knowledge assistants, research automation, multilingual support operations and document-based agents.

Command A pricing and cost trade-offs

The supplied API pricing is $2.50 per 1 million input tokens and $10 per 1 million output tokens. Input tokens are the material sent to the model, while output tokens are the text it generates. Because output is priced at four times the input rate, applications can reduce costs by limiting unnecessary response length, selecting concise instructions and avoiding repeated generation when a shorter result is sufficient.

Command A is not positioned as the lowest-cost option for very high-volume, routine text generation. The editorial cost score is 6 out of 10, reflecting a trade-off between its broad enterprise capabilities and its per-token price. Its value is more apparent when a task benefits from a large context window, tool use, multilingual handling or a reliable structured workflow. A smaller or cheaper model may be more appropriate for simple classification, short summaries or large batches of low-complexity text.

The model has an 8,000-token maximum output. That is enough for substantial reports, extracted records and detailed answers, but it is not unlimited. Applications generating long documents should handle truncation, divide work into stages or use an application-level process that combines multiple responses.

Main strengths and limitations

Strengths

  • Very large context: The 256K-token window supports lengthy source material and broad retrieval contexts.
  • Enterprise workflow fit: Tool use, citations, structured outputs and streaming support agentic applications.
  • Multilingual coverage: Support for 23 languages is useful for organizations serving multiple regions.
  • Strong text reasoning profile: The editorial assessment rates reasoning and coding at 8 out of 10, although these are not provider-published benchmark scores.
  • Production-oriented availability: The model is listed as live and is available through Cohere's API and selected deployment platforms.

Limitations

  • Text only: Command A does not natively accept or generate images, audio or video.
  • No built-in web search: Current information requires retrieval or an external search tool.
  • Knowledge cutoff: The underlying model cutoff is June 1, 2024.
  • Cost: Output pricing is higher than input pricing, and the model may not be economical for simple high-volume generation.
  • Output ceiling: Responses are limited to 8,000 tokens, so very large documents may require multiple stages.
  • Integration requirements: Tool use, retrieval, permissions, citations and validation must be implemented and managed by the surrounding application.

When to choose Command A

Choose Command A when the application needs to combine long-context text processing with enterprise retrieval, tools or structured responses. It is especially suitable when a single workflow must inspect large documents, use external business data, support multiple languages and return information that software can consume.

It is also a reasonable choice for teams building document-focused agents rather than a simple chat interface. The 256K context window can reduce the need to split source material into many small requests, while tool use allows the model to work with systems outside the model itself.

Another model may be more appropriate when the priority is the lowest possible cost, extremely high-volume short-form generation, native image or audio processing, or built-in current web search. A smaller model can also be preferable when the task does not need Command A's long context, multilingual scope or agent features. For current facts, Command A should be paired with a suitable retrieval or search system rather than used alone.

Overall assessment

Cohere Command A is a specialized enterprise language model whose main advantage is the combination of a 256K context window, tool use, structured outputs, multilingual support and agent-oriented design. It is best understood as a foundation for document and workflow applications, not as a consumer multimodal assistant or a standalone source of current information.

Its pricing and text-only design limit its suitability for some workloads, but those trade-offs can be justified when the application needs to process substantial context and connect model reasoning to retrieval or business tools. Teams evaluating it should test their own documents, languages, tool sequences and output schemas, while treating the published limits and pricing as the verified baseline.


Answers to Frequently Asked Questions

What are the main limitations of Cohere Command A?
Command A is a text-only model with no native image, audio or video input or output. Its knowledge cutoff is June 1, 2024, it does not provide built-in web search, and its responses are limited to 8,000 tokens. Retrieval, permissions, tool execution, citation handling and output validation must be managed by the surrounding application.
How much does Cohere Command A cost?
The supplied API pricing is $2.50 per 1 million input tokens and $10.00 per 1 million output tokens. Because output tokens cost four times more than input tokens, applications can manage expenses by limiting unnecessary response length and using smaller models for simple, high-volume tasks.
Does Cohere Command A support RAG, tool use and web search?
Command A supports retrieval-augmented generation and tool use. Applications can connect it to databases, internal systems, calculators or search services. However, it does not have native built-in web search in the supplied specifications; current information must come from an external retrieval or search tool.
What is Cohere Command A?
Cohere Command A is a 111-billion-parameter enterprise language model released on March 13, 2025. It is designed for retrieval-augmented generation, long-context document analysis, multilingual applications, tool-connected assistants, structured text generation and agentic workflows.
What is the context window of Cohere Command A?
Cohere Command A has a 256,000-token context window and supports a maximum output of 8,000 tokens. This allows applications to provide lengthy reports, retrieved passages, policies and other documents in a single request, although the model does not automatically access an organization's data without a retrieval system.


Sources 8
Provider

About Cohere