Command A

Command A Vision

by Cohere · Live

Cohere Command A Vision is a live multimodal Chat API model for enterprise image understanding. It accepts text and images, supports up to 20 images per request, provides a 128K-token context window and 8K-token maximum output, and focuses on OCR, document question answering, chart and table analysis, visual extraction, and multilingual image understanding.

Text Reasoning Coding
Command A Vision extends Cohere’s Command family with image understanding rather than image generation. The model accepts written instructions and visual inputs in the same request, allowing applications to ask questions about documents, charts, tables, diagrams, and other images. Its combination of a large context window, structured extraction support, and enterprise-oriented API access makes it most relevant to document-processing and visual-analysis workflows.
Outputs

What Command A Vision can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming Structured output
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
7/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Command A
Model type Multimodal
Context window 128K tokens
Maximum output 8K tokens
Knowledge cutoff June 1, 2024
Release date 2025-07-31
Status Live
Knowledge cutoff notes

Cohere's model documentation explicitly lists June 1, 2024 as the knowledge cutoff. Image inputs, retrieved context, or application-provided information do not change the underlying cutoff.

Model notes

Canonical model ID is command-a-vision-07-2025. Cohere lists the model as live and supports it through Chat API endpoints. It accepts text and images, supports up to 20 images per request or 20 MB total, and officially supports English, Portuguese, Italian, French, German, and Spanish. Tool use is explicitly unsupported. The model page identifies structured outputs as a capability, but a separate legacy JSON-mode capability was not independently verified for this exact model. Cohere states that the model is free until rate limits are reached and requests that customers contact sales for production use.

Cost

Model pricing

Input Free until applicable rate limits; production access requires contacting Cohere
Output Free until applicable rate limits; production access requires contacting Cohere
Model guide

Command A Vision: Cohere’s Enterprise Model for Document and Image Understanding

Command A Vision is Cohere’s live multimodal model for analyzing images together with text. It is designed for enterprise document intelligence, OCR, chart and table interpretation, visual question answering, and multilingual image understanding, with a 128K-token context window, up to 20 images per request, and text-only output.

What is Command A Vision?

Command A Vision is Cohere’s multimodal model for understanding images alongside natural-language instructions. It is intended for applications that need to extract information from visual material, answer questions about images, or combine visual evidence with text-based context.

The model’s canonical identifier is command-a-vision-07-2025. Cohere lists it as a live model available through its Chat API. Within the Command family, it occupies the image-understanding role: it adds visual input to the family’s text-generation interface while keeping the output modality text.

This distinction matters. Command A Vision can inspect and describe visual content, but it does not create images. Its practical value is in turning visual information into answers, explanations, classifications, or structured text that another application can process.

What can Command A Vision analyze?

Cohere documents support for text and image inputs in one request. Images may be supplied through HTTP image URLs or base64 data URLs. Up to 20 images can be included in a request, subject to a total image limit of 20 MB.

Typical tasks include:

  • Document question answering: Ask questions about forms, reports, scanned pages, screenshots, or other visual documents.
  • OCR and extraction: Read text from images and convert relevant content into a usable textual or structured representation.
  • Charts and graphs: Identify trends, compare categories, summarize plotted information, or answer questions about chart values.
  • Tables: Extract selected rows or columns, compare values, and interpret tables embedded in images.
  • Diagrams and visual layouts: Interpret relationships and labels in diagrams or other structured visual material.
  • Object and scene analysis: Describe or identify visible objects and relevant visual details.
  • Multilingual visual understanding: Process supported languages appearing in visual content or used in the accompanying prompt.

These capabilities are especially useful when the source information is not already available as clean machine-readable text. For example, an application could submit several pages of a scanned report and ask for specific figures, differences between pages, or a structured summary of the findings.

Technical specifications and limits

SpecificationVerified detail
ProviderCohere
Model IDcommand-a-vision-07-2025
StatusLive
Context window128,000 tokens
Maximum output8,000 tokens
Input modalitiesText and images
Output modalityText
Image limitUp to 20 images per request and 20 MB total
Knowledge cutoffJune 1, 2024
API interfaceChat API
Tool useNot supported

The 128,000-token context window describes how much text and model context can be handled in a request; it does not remove the separate image-count and image-size limits. The maximum generated response is 8,000 tokens. Actual usage, latency, and cost can still vary with prompt length, image detail, the number of images, and the requested response size.

Cohere identifies English, Portuguese, Italian, French, German, and Spanish as supported languages for Command A Vision. Results may vary according to image resolution, typography, layout, visual complexity, and how clearly the question is phrased.

Strengths for enterprise visual workflows

Document intelligence and OCR

Command A Vision is a practical fit when an organization receives information as scans, screenshots, photographed pages, or other visual documents. It can answer questions about those sources and help convert their contents into text or structured fields. This can support workflows such as form review, report extraction, multilingual document processing, and visual records search.

The model is not limited to reading isolated text. Its intended use also includes understanding the relationship between text, layout, tables, and other visual elements. That makes it more suitable for document questions where location and presentation affect meaning.

Charts, tables, and diagrams

Many extraction systems work well with plain text but lose information when data is embedded in a chart or table image. Command A Vision is designed to interpret those formats. A prompt might ask it to identify the highest category, compare two periods, extract a particular row, or summarize a trend.

For production workflows, applications should still validate extracted values when accuracy is important. Small labels, low-resolution images, unusual chart designs, and dense tables can make visual interpretation more difficult.

Large context and multiple images

The 128K-token context window gives applications room to combine substantial textual instructions or supporting material with image inputs. Support for up to 20 images per request is also useful for multi-page documents, before-and-after comparisons, and sets of related figures. The total 20 MB image limit remains a practical constraint, so large source files may need to be resized, split, or otherwise prepared before submission.

Reasoning, coding, and tool support

Command A Vision is designed for visual interpretation and text generation rather than autonomous action. Cohere’s supplied model assessment rates its reasoning and coding capabilities at 6 out of 10, but these are editorial scores in the supplied data, not provider-published benchmark results. They should be treated as directional evaluations rather than formal measurements.

In practical terms, the model can reason over information visible in an image: it can compare values, follow a visual question, summarize evidence, and return requested fields. It can also generate code or structured text when prompted, but the research does not establish it as a specialized coding model.

Tool use is explicitly unsupported for this model. Command A Vision therefore cannot independently call external functions, browse the web, or execute a workflow through native tool calling. If an application needs current information or an external action, the surrounding software must retrieve the information or perform the action separately and then provide relevant context to the model.

Cohere identifies structured outputs as a capability. However, a separate legacy JSON-mode capability was not independently verified for this exact model. Developers should distinguish structured-output support from assuming that every JSON-specific mode or SDK option is available.

Pricing and API access

Cohere’s model documentation states that Command A Vision is free until applicable rate limits are reached and directs customers to contact Cohere for production use. The supplied research does not provide a published per-token price for the model, so there is no verified recurring or per-token price to report.

Cohere lists a trial limit of 20 requests per minute for Command A Vision. Trial access is therefore useful for evaluation and prototyping, but it should not be treated as an unlimited production allocation. Production availability requires contacting Cohere and may involve account, billing, capacity, and deployment arrangements that are not specified in the supplied model documentation.

The model uses Cohere’s Chat API interaction pattern. Image inputs can be sent by URL or as base64 data URLs, and the response is text. Applications should account for image preparation, request size, rate limits, and response length when estimating operational cost and latency.

Limitations and trade-offs

  • No image generation: Command A Vision understands images but does not produce image outputs.
  • No native tool use: It cannot independently call tools or take external actions.
  • Text output only: The model returns textual results rather than edited images, audio, or video.
  • Image constraints: Requests are limited to 20 images and 20 MB total, even though the text context window is much larger.
  • Knowledge cutoff: Its stated cutoff is June 1, 2024. Current facts must be supplied by the application or obtained through a separate retrieval system.
  • Visual quality dependence: Resolution, legibility, typography, layout, and image complexity can affect results.
  • Production pricing is not public in the supplied research: Trial access is rate-limited, while production use requires contacting Cohere.

These constraints make the model less appropriate for an image-generation pipeline, an agent that must browse and operate software, or a workflow requiring built-in access to current web information. A model with native tool calling may be more suitable for those tasks, while a dedicated image-generation system is the appropriate choice when the required output is a new image.

When to choose Command A Vision

Choose Command A Vision when the central problem is understanding visual information and turning it into useful text. It is particularly well suited to enterprise applications that need document question answering, OCR, chart and table analysis, visual extraction, or multilingual processing through a Cohere-managed API.

It is a strong candidate when an application must work with multiple related images, such as pages in a scanned document or a sequence of charts, and when a large textual context is useful for instructions or supporting material. Its text-only output can also fit downstream systems that store extracted fields, summaries, classifications, or answers.

Consider another option when image creation, audio or video processing, autonomous tool calls, web search, or highly specialized coding is the primary requirement. Also consider a different deployment or model choice if your workflow needs a publicly documented production price rather than a sales-led arrangement. For current information, pair Command A Vision with an external retrieval layer and clearly pass the retrieved evidence in the prompt.

Overall assessment

Command A Vision is a focused enterprise vision-language model rather than a general consumer assistant. Its verified profile is centered on text-and-image input, a 128K-token context window, 8K-token maximum output, support for up to 20 images per request, and text responses through Cohere’s Chat API.

Its clearest value is visual document and data understanding: reading, comparing, extracting, and explaining information contained in images. The absence of image generation and native tools narrows its scope, but those limitations are consistent with its role. For organizations building document intelligence or visual-analysis workflows, it offers a dedicated model option; for action-oriented or generative media tasks, a different type of system will be more appropriate.


Answers to Frequently Asked Questions

When should an organization choose Command A Vision?
Organizations should choose Command A Vision when they need to understand scanned documents, screenshots, charts, tables, diagrams, or other visual content and convert that information into useful text. It is less suitable for image generation, autonomous tool use, web search, audio or video processing, or highly specialized coding.
How can developers access Command A Vision and what does it cost?
Command A Vision is available through Cohere’s Chat API using the model ID "command-a-vision-07-2025." Cohere states that it is free until applicable rate limits are reached, with a trial limit of 20 requests per minute. No verified public per-token production price is provided; production users should contact Cohere.
What are the image and context limits of Command A Vision?
Command A Vision supports up to 20 images per request with a total image size limit of 20 MB. It has a 128,000-token context window and can generate up to 8,000 tokens of text output.
Can Command A Vision generate images or use external tools?
No. Command A Vision analyzes images and returns text, but it does not generate images. Native tool use is also unsupported, so applications must handle web retrieval, external function calls, and other actions separately.
What is Command A Vision used for?
Command A Vision is Cohere’s multimodal model for understanding images alongside natural-language instructions. It is designed for document question answering, OCR and data extraction, chart and table analysis, diagram interpretation, object recognition, and converting visual information into text or structured outputs.


Sources 5
Provider

About Cohere