What is Aya Vision 32B?
Aya Vision 32B is Cohere Labs' 32-billion-parameter multilingual vision-language model. In practical terms, it can read text and inspect images before producing a text answer. That makes it suitable for tasks such as describing a photograph, extracting writing from a document, answering questions about a chart, translating text found in an image, or summarizing visual content.
The model's official Cohere API identifier is c4ai-aya-vision-32b. The downloadable model is published on Hugging Face as CohereLabs/aya-vision-32b. It belongs to the Aya Vision family and is positioned for multilingual image understanding rather than media generation.
Aya Vision 32B is not an image-generation, speech, music or video model. Its output is text, even when the input includes one or more images.
Supported languages and modalities
The model accepts text and images and returns text. Cohere's documentation describes support for 23 languages: English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Arabic, Simplified Chinese, Traditional Chinese, Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew and Persian.
This language coverage is important because the model is designed for visual-language work outside English-only workflows. For example, a user could submit an image containing Arabic or Japanese text and ask for transcription, translation or a summary in another supported language. The supplied documentation supports these use cases, but it does not establish that every language performs equally well on every task.
| Capability | Supported behavior |
|---|---|
| Text input | Yes |
| Image input | Yes |
| Text output | Yes |
| Image, video or audio output | No |
| Document and image understanding | Yes |
| Web search | Not documented for this model |
What can Aya Vision 32B do?
Aya Vision 32B is intended for applications where visual information must be interpreted alongside language. Its documented and practical use cases include:
- Optical character recognition: reading printed or handwritten-looking text in images and documents, where image quality permits.
- Image captioning: generating descriptions of scenes, objects and visual content.
- Visual question answering: answering questions about the contents, relationships or apparent meaning of an image.
- Image-based translation: translating text shown in signs, screenshots, forms or other visual material.
- Classification and extraction: sorting or identifying visual content and extracting useful information from it.
- Summarization: turning visual documents or image-based information into a shorter text explanation.
- Multilingual visual reasoning: combining image interpretation with reasoning and responses in supported languages.
These capabilities make the model relevant to accessibility tools, multilingual document processing, research prototypes, visual search components and workflows that need to inspect images before generating written output.
Architecture, context and processing limits
The model combines the Aya Expanse 32B language model with a SigLIP2-patch14-384 vision encoder and a multimodal adapter. The language model handles the text-generation portion, while the vision encoder converts image content into information the language model can use.
Aya Vision 32B has approximately 32 billion parameters. Cohere documents a 16K-token context window and a maximum output of 4K tokens when the model is used through its API. The context window includes the text and visual information supplied to the model, so long prompts, multiple images or image-heavy documents can reduce the space available for the response.
For image processing, the model represents images using tiles at a 364-by-364-pixel resolution. Documentation describes support for up to 12 image tiles plus a thumbnail, with up to 2,197 image tokens for this representation. Images with different aspect ratios are mapped to supported resolutions. These details matter for applications involving large, detailed or unusually shaped images: resizing and tiling can affect both the information presented to the model and the amount of context consumed.
Access, pricing and license
Aya Vision 32B can be accessed through Cohere's Chat API with the model ID c4ai-aya-vision-32b. It is also available as downloadable open weights for compatible Transformers, vLLM, SGLang, Docker and local inference workflows.
Cohere's public pricing information does not publish a model-specific price for Aya Vision 32B. The pricing page lists pricing for some other Aya models, including Aya Expanse, but that should not be treated as the price of Aya Vision 32B. API users should verify current account-specific pricing and availability with Cohere before estimating operating costs.
The downloadable weights use the CC-BY-NC-4.0 license and also require compliance with Cohere Labs' Acceptable Use Policy. The non-commercial license is a significant limitation for companies planning to deploy the weights in a commercial product. Commercial users should review the current license and applicable Cohere terms rather than assuming that open weights permit unrestricted commercial use.
Strengths and trade-offs
The clearest strength of Aya Vision 32B is its combination of visual understanding and broad multilingual coverage. A model that can inspect images and respond in many languages can reduce the need to build separate OCR, translation and captioning components for every supported language. Its open-weight availability also gives research and infrastructure teams more deployment flexibility than a model available only through a hosted endpoint.
That flexibility comes with hardware and operational costs. A 32-billion-parameter model is substantially heavier to run than a compact vision-language model, particularly when using the original unquantized weights. Local deployment can therefore require more memory, specialized infrastructure and optimization work. Hosted API access may be simpler, but its model-specific pricing and production terms should be confirmed before deployment.
The model also generates text only. Teams seeking image creation, video generation, speech synthesis or audio understanding need a different model or a separate service. The supplied documentation does not establish a dedicated web-search feature, structured-output guarantee, prompt caching, batch API or model-specific knowledge-cutoff date.
Reasoning, coding and tool support
Aya Vision 32B can perform visual reasoning in the ordinary sense of interpreting an image, connecting visual evidence to a question and explaining the result. However, the supplied research does not identify a separate reasoning mode or provider-published reasoning benchmark for this model.
The editorial dataset rates its reasoning at 7 out of 10 and coding at 6 out of 10. These are comparative editorial estimates, not Cohere-published scores and not standardized benchmark results. Coding may be useful when the model is asked to write or explain code based on a visual input, but Aya Vision 32B is primarily a visual-language model rather than a specialist programming model.
Tool use and function calling are not verified in the supplied model-specific research. The model should therefore not be selected on the assumption that it can browse the web, execute code, call external functions or operate business systems without an additional, documented integration layer.
When to choose Aya Vision 32B
Aya Vision 32B is a strong candidate when the central problem is multilingual image understanding and the team values access to open weights. It is especially suitable for:
- OCR and translation across a wide range of supported languages.
- Image captioning and accessibility-oriented descriptions.
- Visual question answering over images, screenshots and documents.
- Research or private deployments that need more control than a hosted-only model provides.
- Applications that can tolerate the infrastructure requirements of a 32-billion-parameter model.
A hosted vision-language model may be more appropriate when rapid deployment, predictable managed infrastructure or lower operational complexity matters more than downloadable weights. A smaller vision-language model may be preferable for low-latency applications, limited hardware or high-volume processing where the larger model's capability does not justify its resource requirements.
Another option is also preferable when the required output is an image, video, audio file or speech signal rather than text. Likewise, teams that require guaranteed JSON schemas, built-in web search, code execution, function calling or a documented batch workflow should verify those capabilities separately instead of assuming they are included.
Bottom line
Aya Vision 32B is best understood as a multilingual image-understanding model with text output, not as a general media-generation system. Its 23-language coverage, image capabilities and downloadable weights make it useful for OCR, translation, captioning and visual question answering. The main trade-offs are its substantial hardware requirements, the lack of a published model-specific API price in the supplied information, and the non-commercial license attached to the downloadable weights.

