What is Command A Vision?
Command A Vision is Cohere’s multimodal model for understanding images alongside natural-language instructions. It is intended for applications that need to extract information from visual material, answer questions about images, or combine visual evidence with text-based context.
The model’s canonical identifier is command-a-vision-07-2025. Cohere lists it as a live model available through its Chat API. Within the Command family, it occupies the image-understanding role: it adds visual input to the family’s text-generation interface while keeping the output modality text.
This distinction matters. Command A Vision can inspect and describe visual content, but it does not create images. Its practical value is in turning visual information into answers, explanations, classifications, or structured text that another application can process.
What can Command A Vision analyze?
Cohere documents support for text and image inputs in one request. Images may be supplied through HTTP image URLs or base64 data URLs. Up to 20 images can be included in a request, subject to a total image limit of 20 MB.
Typical tasks include:
- Document question answering: Ask questions about forms, reports, scanned pages, screenshots, or other visual documents.
- OCR and extraction: Read text from images and convert relevant content into a usable textual or structured representation.
- Charts and graphs: Identify trends, compare categories, summarize plotted information, or answer questions about chart values.
- Tables: Extract selected rows or columns, compare values, and interpret tables embedded in images.
- Diagrams and visual layouts: Interpret relationships and labels in diagrams or other structured visual material.
- Object and scene analysis: Describe or identify visible objects and relevant visual details.
- Multilingual visual understanding: Process supported languages appearing in visual content or used in the accompanying prompt.
These capabilities are especially useful when the source information is not already available as clean machine-readable text. For example, an application could submit several pages of a scanned report and ask for specific figures, differences between pages, or a structured summary of the findings.
Technical specifications and limits
| Specification | Verified detail |
|---|---|
| Provider | Cohere |
| Model ID | command-a-vision-07-2025 |
| Status | Live |
| Context window | 128,000 tokens |
| Maximum output | 8,000 tokens |
| Input modalities | Text and images |
| Output modality | Text |
| Image limit | Up to 20 images per request and 20 MB total |
| Knowledge cutoff | June 1, 2024 |
| API interface | Chat API |
| Tool use | Not supported |
The 128,000-token context window describes how much text and model context can be handled in a request; it does not remove the separate image-count and image-size limits. The maximum generated response is 8,000 tokens. Actual usage, latency, and cost can still vary with prompt length, image detail, the number of images, and the requested response size.
Cohere identifies English, Portuguese, Italian, French, German, and Spanish as supported languages for Command A Vision. Results may vary according to image resolution, typography, layout, visual complexity, and how clearly the question is phrased.
Strengths for enterprise visual workflows
Document intelligence and OCR
Command A Vision is a practical fit when an organization receives information as scans, screenshots, photographed pages, or other visual documents. It can answer questions about those sources and help convert their contents into text or structured fields. This can support workflows such as form review, report extraction, multilingual document processing, and visual records search.
The model is not limited to reading isolated text. Its intended use also includes understanding the relationship between text, layout, tables, and other visual elements. That makes it more suitable for document questions where location and presentation affect meaning.
Charts, tables, and diagrams
Many extraction systems work well with plain text but lose information when data is embedded in a chart or table image. Command A Vision is designed to interpret those formats. A prompt might ask it to identify the highest category, compare two periods, extract a particular row, or summarize a trend.
For production workflows, applications should still validate extracted values when accuracy is important. Small labels, low-resolution images, unusual chart designs, and dense tables can make visual interpretation more difficult.
Large context and multiple images
The 128K-token context window gives applications room to combine substantial textual instructions or supporting material with image inputs. Support for up to 20 images per request is also useful for multi-page documents, before-and-after comparisons, and sets of related figures. The total 20 MB image limit remains a practical constraint, so large source files may need to be resized, split, or otherwise prepared before submission.
Reasoning, coding, and tool support
Command A Vision is designed for visual interpretation and text generation rather than autonomous action. Cohere’s supplied model assessment rates its reasoning and coding capabilities at 6 out of 10, but these are editorial scores in the supplied data, not provider-published benchmark results. They should be treated as directional evaluations rather than formal measurements.
In practical terms, the model can reason over information visible in an image: it can compare values, follow a visual question, summarize evidence, and return requested fields. It can also generate code or structured text when prompted, but the research does not establish it as a specialized coding model.
Tool use is explicitly unsupported for this model. Command A Vision therefore cannot independently call external functions, browse the web, or execute a workflow through native tool calling. If an application needs current information or an external action, the surrounding software must retrieve the information or perform the action separately and then provide relevant context to the model.
Cohere identifies structured outputs as a capability. However, a separate legacy JSON-mode capability was not independently verified for this exact model. Developers should distinguish structured-output support from assuming that every JSON-specific mode or SDK option is available.
Pricing and API access
Cohere’s model documentation states that Command A Vision is free until applicable rate limits are reached and directs customers to contact Cohere for production use. The supplied research does not provide a published per-token price for the model, so there is no verified recurring or per-token price to report.
Cohere lists a trial limit of 20 requests per minute for Command A Vision. Trial access is therefore useful for evaluation and prototyping, but it should not be treated as an unlimited production allocation. Production availability requires contacting Cohere and may involve account, billing, capacity, and deployment arrangements that are not specified in the supplied model documentation.
The model uses Cohere’s Chat API interaction pattern. Image inputs can be sent by URL or as base64 data URLs, and the response is text. Applications should account for image preparation, request size, rate limits, and response length when estimating operational cost and latency.
Limitations and trade-offs
- No image generation: Command A Vision understands images but does not produce image outputs.
- No native tool use: It cannot independently call tools or take external actions.
- Text output only: The model returns textual results rather than edited images, audio, or video.
- Image constraints: Requests are limited to 20 images and 20 MB total, even though the text context window is much larger.
- Knowledge cutoff: Its stated cutoff is June 1, 2024. Current facts must be supplied by the application or obtained through a separate retrieval system.
- Visual quality dependence: Resolution, legibility, typography, layout, and image complexity can affect results.
- Production pricing is not public in the supplied research: Trial access is rate-limited, while production use requires contacting Cohere.
These constraints make the model less appropriate for an image-generation pipeline, an agent that must browse and operate software, or a workflow requiring built-in access to current web information. A model with native tool calling may be more suitable for those tasks, while a dedicated image-generation system is the appropriate choice when the required output is a new image.
When to choose Command A Vision
Choose Command A Vision when the central problem is understanding visual information and turning it into useful text. It is particularly well suited to enterprise applications that need document question answering, OCR, chart and table analysis, visual extraction, or multilingual processing through a Cohere-managed API.
It is a strong candidate when an application must work with multiple related images, such as pages in a scanned document or a sequence of charts, and when a large textual context is useful for instructions or supporting material. Its text-only output can also fit downstream systems that store extracted fields, summaries, classifications, or answers.
Consider another option when image creation, audio or video processing, autonomous tool calls, web search, or highly specialized coding is the primary requirement. Also consider a different deployment or model choice if your workflow needs a publicly documented production price rather than a sales-led arrangement. For current information, pair Command A Vision with an external retrieval layer and clearly pass the retrieved evidence in the prompt.
Overall assessment
Command A Vision is a focused enterprise vision-language model rather than a general consumer assistant. Its verified profile is centered on text-and-image input, a 128K-token context window, 8K-token maximum output, support for up to 20 images per request, and text responses through Cohere’s Chat API.
Its clearest value is visual document and data understanding: reading, comparing, extracting, and explaining information contained in images. The absence of image generation and native tools narrows its scope, but those limitations are consistent with its role. For organizations building document intelligence or visual-analysis workflows, it offers a dedicated model option; for action-oriented or generative media tasks, a different type of system will be more appropriate.

