Aya Vision

Aya Vision 32B

by Cohere · Live

Cohere Labs' Aya Vision 32B is an open-weights multilingual vision-language model that accepts text and images and generates text in 23 languages. It offers a 16K context window, up to 4K API output tokens and support for OCR, captioning, image translation, classification and visual question answering. The article covers its architecture, access options, licensing, limitations and practical trade-offs.

Text Reasoning Coding
Aya Vision 32B is a multimodal model from Cohere Labs that accepts text and images and responds with text. It combines the Aya Expanse 32B language model with a SigLIP2 vision encoder, supports a 16K-token context window and can produce up to 4K output tokens through Cohere's API. The model is available through the Cohere Chat API and as downloadable open weights, although the downloadable model uses a non-commercial license.
Outputs

What Aya Vision 32B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

7/10 Reasoning
6/10 Coding
5/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Aya Vision
Model type Multimodal
Context window 16K tokens
Maximum output 4K tokens
Release date 2025-03-04
Status Live
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was found in Cohere's documentation or the official model card.

Model notes

Aya Vision 32B is the 32-billion-parameter member of the Aya Vision family. The Cohere API identifier is c4ai-aya-vision-32b and the open-weight Hugging Face identifier is CohereLabs/aya-vision-32b. It accepts text and images and returns text. Cohere documents a 16K context window and 4K maximum output through the API. The model covers 23 languages and uses Aya Expanse 32B with a SigLIP2-patch14-384 vision encoder and multimodal adapter. Cohere's public pricing page lists pricing for Aya Expanse but does not publish a model-specific price for Aya Vision 32B. The downloadable weights use CC-BY-NC-4.0 and require compliance with Cohere Labs' Acceptable Use Policy. Editorial scores are comparative estimates, not provider-published benchmarks.

Model guide

Aya Vision 32B: Multilingual Image Understanding with Open Weights

Cohere Labs' Aya Vision 32B is a 32-billion-parameter open-weights vision-language model for understanding images and text across 23 languages. It generates text rather than images, and is aimed at OCR, captioning, visual question answering, image translation, classification, summarization and multilingual visual reasoning.

What is Aya Vision 32B?

Aya Vision 32B is Cohere Labs' 32-billion-parameter multilingual vision-language model. In practical terms, it can read text and inspect images before producing a text answer. That makes it suitable for tasks such as describing a photograph, extracting writing from a document, answering questions about a chart, translating text found in an image, or summarizing visual content.

The model's official Cohere API identifier is c4ai-aya-vision-32b. The downloadable model is published on Hugging Face as CohereLabs/aya-vision-32b. It belongs to the Aya Vision family and is positioned for multilingual image understanding rather than media generation.

Aya Vision 32B is not an image-generation, speech, music or video model. Its output is text, even when the input includes one or more images.

Supported languages and modalities

The model accepts text and images and returns text. Cohere's documentation describes support for 23 languages: English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Arabic, Simplified Chinese, Traditional Chinese, Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew and Persian.

This language coverage is important because the model is designed for visual-language work outside English-only workflows. For example, a user could submit an image containing Arabic or Japanese text and ask for transcription, translation or a summary in another supported language. The supplied documentation supports these use cases, but it does not establish that every language performs equally well on every task.

CapabilitySupported behavior
Text inputYes
Image inputYes
Text outputYes
Image, video or audio outputNo
Document and image understandingYes
Web searchNot documented for this model

What can Aya Vision 32B do?

Aya Vision 32B is intended for applications where visual information must be interpreted alongside language. Its documented and practical use cases include:

  • Optical character recognition: reading printed or handwritten-looking text in images and documents, where image quality permits.
  • Image captioning: generating descriptions of scenes, objects and visual content.
  • Visual question answering: answering questions about the contents, relationships or apparent meaning of an image.
  • Image-based translation: translating text shown in signs, screenshots, forms or other visual material.
  • Classification and extraction: sorting or identifying visual content and extracting useful information from it.
  • Summarization: turning visual documents or image-based information into a shorter text explanation.
  • Multilingual visual reasoning: combining image interpretation with reasoning and responses in supported languages.

These capabilities make the model relevant to accessibility tools, multilingual document processing, research prototypes, visual search components and workflows that need to inspect images before generating written output.

Architecture, context and processing limits

The model combines the Aya Expanse 32B language model with a SigLIP2-patch14-384 vision encoder and a multimodal adapter. The language model handles the text-generation portion, while the vision encoder converts image content into information the language model can use.

Aya Vision 32B has approximately 32 billion parameters. Cohere documents a 16K-token context window and a maximum output of 4K tokens when the model is used through its API. The context window includes the text and visual information supplied to the model, so long prompts, multiple images or image-heavy documents can reduce the space available for the response.

For image processing, the model represents images using tiles at a 364-by-364-pixel resolution. Documentation describes support for up to 12 image tiles plus a thumbnail, with up to 2,197 image tokens for this representation. Images with different aspect ratios are mapped to supported resolutions. These details matter for applications involving large, detailed or unusually shaped images: resizing and tiling can affect both the information presented to the model and the amount of context consumed.

Access, pricing and license

Aya Vision 32B can be accessed through Cohere's Chat API with the model ID c4ai-aya-vision-32b. It is also available as downloadable open weights for compatible Transformers, vLLM, SGLang, Docker and local inference workflows.

Cohere's public pricing information does not publish a model-specific price for Aya Vision 32B. The pricing page lists pricing for some other Aya models, including Aya Expanse, but that should not be treated as the price of Aya Vision 32B. API users should verify current account-specific pricing and availability with Cohere before estimating operating costs.

The downloadable weights use the CC-BY-NC-4.0 license and also require compliance with Cohere Labs' Acceptable Use Policy. The non-commercial license is a significant limitation for companies planning to deploy the weights in a commercial product. Commercial users should review the current license and applicable Cohere terms rather than assuming that open weights permit unrestricted commercial use.

Strengths and trade-offs

The clearest strength of Aya Vision 32B is its combination of visual understanding and broad multilingual coverage. A model that can inspect images and respond in many languages can reduce the need to build separate OCR, translation and captioning components for every supported language. Its open-weight availability also gives research and infrastructure teams more deployment flexibility than a model available only through a hosted endpoint.

That flexibility comes with hardware and operational costs. A 32-billion-parameter model is substantially heavier to run than a compact vision-language model, particularly when using the original unquantized weights. Local deployment can therefore require more memory, specialized infrastructure and optimization work. Hosted API access may be simpler, but its model-specific pricing and production terms should be confirmed before deployment.

The model also generates text only. Teams seeking image creation, video generation, speech synthesis or audio understanding need a different model or a separate service. The supplied documentation does not establish a dedicated web-search feature, structured-output guarantee, prompt caching, batch API or model-specific knowledge-cutoff date.

Reasoning, coding and tool support

Aya Vision 32B can perform visual reasoning in the ordinary sense of interpreting an image, connecting visual evidence to a question and explaining the result. However, the supplied research does not identify a separate reasoning mode or provider-published reasoning benchmark for this model.

The editorial dataset rates its reasoning at 7 out of 10 and coding at 6 out of 10. These are comparative editorial estimates, not Cohere-published scores and not standardized benchmark results. Coding may be useful when the model is asked to write or explain code based on a visual input, but Aya Vision 32B is primarily a visual-language model rather than a specialist programming model.

Tool use and function calling are not verified in the supplied model-specific research. The model should therefore not be selected on the assumption that it can browse the web, execute code, call external functions or operate business systems without an additional, documented integration layer.

When to choose Aya Vision 32B

Aya Vision 32B is a strong candidate when the central problem is multilingual image understanding and the team values access to open weights. It is especially suitable for:

  • OCR and translation across a wide range of supported languages.
  • Image captioning and accessibility-oriented descriptions.
  • Visual question answering over images, screenshots and documents.
  • Research or private deployments that need more control than a hosted-only model provides.
  • Applications that can tolerate the infrastructure requirements of a 32-billion-parameter model.

A hosted vision-language model may be more appropriate when rapid deployment, predictable managed infrastructure or lower operational complexity matters more than downloadable weights. A smaller vision-language model may be preferable for low-latency applications, limited hardware or high-volume processing where the larger model's capability does not justify its resource requirements.

Another option is also preferable when the required output is an image, video, audio file or speech signal rather than text. Likewise, teams that require guaranteed JSON schemas, built-in web search, code execution, function calling or a documented batch workflow should verify those capabilities separately instead of assuming they are included.

Bottom line

Aya Vision 32B is best understood as a multilingual image-understanding model with text output, not as a general media-generation system. Its 23-language coverage, image capabilities and downloadable weights make it useful for OCR, translation, captioning and visual question answering. The main trade-offs are its substantial hardware requirements, the lack of a published model-specific API price in the supplied information, and the non-commercial license attached to the downloadable weights.


Answers to Frequently Asked Questions

What license applies to the Aya Vision 32B open weights?
The downloadable Aya Vision 32B weights are provided under the CC-BY-NC-4.0 license and require compliance with Cohere Labs' Acceptable Use Policy. The non-commercial license may prevent use in commercial products, so organizations should review the current license and Cohere terms before deployment.
How can developers access Aya Vision 32B?
Developers can access it through Cohere's Chat API using the model ID "c4ai-aya-vision-32b" or download the open weights from Hugging Face as "CohereLabs/aya-vision-32b" for compatible Transformers, vLLM, SGLang, Docker and local inference workflows.
Can Aya Vision 32B generate images, video or audio?
No. Aya Vision 32B accepts text and images but produces text only. It is designed for image understanding rather than image, video, audio or speech generation.
What is Aya Vision 32B used for?
Aya Vision 32B is a multilingual vision-language model used for OCR, image captioning, visual question answering, image-based translation, document understanding, classification, information extraction and summarizing visual content.
Which languages does Aya Vision 32B support?
Aya Vision 32B supports 23 languages, including English, French, Spanish, German, Portuguese, Japanese, Korean, Arabic, Simplified and Traditional Chinese, Russian, Turkish, Vietnamese, Hindi, Hebrew and Persian.


Sources 7
Provider

About Cohere