What is Aya Expanse 32B?
Aya Expanse 32B is a multilingual large language model from Cohere Labs, the research organization associated with Cohere. The model has 32 billion parameters and is designed primarily for text generation rather than image, audio, or video processing. Its canonical Cohere API identifier is c4ai-aya-expanse-32b, and its open-weight release is available through the gated Hugging Face repository CohereLabs/aya-expanse-32b.
The model is aimed at applications where language coverage matters as much as general text quality. Typical uses include translation assistance, multilingual writing, summarization, cross-language customer support, international communications, educational tools, and research involving multilingual language models. It can also be used for long-document work because its documented context window is 128,000 tokens.
Aya Expanse 32B fits within Cohere’s broader model lineup as a multilingual generation model. It is distinct from Cohere’s retrieval-focused Embed and Rerank families, document-processing tools such as Parse, and the company’s enterprise workspace products. Its primary role is to generate and transform text across many languages.
Languages and multilingual design
Aya Expanse 32B supports 23 languages: Arabic; simplified and traditional Chinese; Czech; Dutch; English; French; German; Greek; Hebrew; Hindi; Indonesian; Italian; Japanese; Korean; Persian; Polish; Portuguese; Romanian; Russian; Spanish; Turkish; Ukrainian; and Vietnamese.
Its main distinction is this broad multilingual focus, not multimodal input or specialized reasoning. Cohere Labs describes a training approach that combines supervised fine-tuning, multilingual preference training, safety tuning, data-arbitrage and data-selection techniques, and model merging. In practical terms, this is intended to help the model follow instructions and produce useful text across languages instead of concentrating most of its quality on English.
For example, a support application could use Aya Expanse 32B to summarize a customer request in one language and draft a response in another. A research workflow could use it to translate source material, create summaries, or compare information across languages. These are suitable applications because they use the model’s central strength: language transformation across a broad set of supported languages.
Context window and output limits
The model has a documented context window of 128,000 tokens. A context window is the amount of input and conversation history the model can consider in one request, measured in tokens rather than words. The large limit makes Aya Expanse 32B suitable for long documents, extended conversations, and multilingual material that would exceed the limits of smaller-context models.
Cohere documents a maximum API output of 4,000 tokens for the Chat API. The context limit and output limit serve different purposes: the 128K window covers the material supplied to the model, while the 4K maximum controls how much text it can generate in the response. The supplied documentation does not specify a knowledge-cutoff date for Aya Expanse 32B, so users should not assume that it contains current factual information.
API and open-weight deployment
There are two main ways to use Aya Expanse 32B. Developers can call it through Cohere’s Chat API using the model identifier c4ai-aya-expanse-32b. This avoids managing model files and hardware, but usage is billed according to token consumption and depends on the API arrangement and service availability.
The second option is the gated open-weight release on Hugging Face. The repository is intended for research and non-commercial use under the CC-BY-NC-4.0 license, together with Cohere Labs’ acceptable-use requirements. Access, licensing, and the intended deployment scenario should be reviewed before downloading or serving the model.
The open-weight model can be loaded with the Transformers ecosystem and served with tools such as vLLM or SGLang. A full-precision 32-billion-parameter model requires substantial GPU memory and inference capacity. Quantized community versions can reduce memory requirements, although quantization may reduce output quality. The supplied research does not establish a single hardware requirement, so actual needs will depend on precision, serving configuration, concurrency, and the chosen implementation.
API pricing
Cohere’s current pricing information lists Aya Expanse models, including Aya Expanse 32B, at $0.50 per million input tokens and $1.50 per million output tokens. Input tokens are the text sent to the model, while output tokens are the text it generates.
This is hosted API pricing, not a price for downloading or operating the open-weight release. Self-hosting introduces infrastructure, storage, engineering, and maintenance costs, even when the model weights themselves are available under the non-commercial license. Commercial users should verify whether the Cohere API or another authorized hosting arrangement permits their intended use; the open-weight CC-BY-NC-4.0 terms do not generally provide unrestricted commercial self-hosting rights.
Capabilities and supported modalities
Aya Expanse 32B is a text-in, text-out model. It accepts text prompts and generates text responses. The research identifies no native image, audio, or video input, and no image, audio, video, music, speech, or embedding output. It should therefore not be selected as a direct multimodal model for interpreting photographs, transcribing recordings, analyzing video, or generating media.
Its strongest supported tasks are multilingual text generation, translation, summarization, customer-support drafting, content creation, and global communication. The model can be incorporated into larger systems that provide document retrieval or other preprocessing, but those surrounding functions should not be confused with capabilities native to Aya Expanse 32B itself.
Tool use, function calling, streaming, caching, batch API support, and a distinct JSON mode are not established by the supplied research. Developers should not assume that these features are available simply because the model can be accessed through a Chat API. Structured application output may require application-side validation or a separately documented provider feature.
Reasoning, coding, speed, and cost trade-offs
Aya Expanse 32B is a general-purpose multilingual generator rather than a model specifically optimized for advanced reasoning or software engineering. It can produce explanations, transform text, and assist with ordinary language-oriented coding prompts, but the supplied research does not establish specialized coding performance or advanced reasoning benchmarks.
The model’s editorial reasoning score is 6 out of 10 and coding score is 5 out of 10. These are comparative editorial estimates, not provider-published benchmark results. They should be treated as directional judgments: Aya Expanse 32B is more compelling for multilingual language work than for demanding mathematical reasoning, complex agent planning, or code-heavy workflows.
Its editorial speed score is 4 out of 10 and cost score is 7 out of 10. These scores are also comparative estimates rather than verified provider metrics. The practical trade-off is straightforward: a 32-billion-parameter model with a 128K context window can provide substantial capacity for multilingual and long-context tasks, but it may require more compute and respond more slowly than smaller models. API users pay relatively low token rates in the supplied pricing information, while self-hosting users must account for GPU and operational costs.
Main strengths and limitations
Strengths
- Broad language coverage: it supports 23 languages, making it useful for cross-language workflows and international applications.
- Long context: the 128K context window supports large documents and extended interactions.
- Open-weight availability: researchers and eligible non-commercial users can examine and deploy the gated release rather than relying exclusively on a hosted endpoint.
- Useful text tasks: translation, summarization, multilingual writing, customer support, and international communication align closely with its intended role.
- Flexible deployment choices: users can select Cohere’s API or investigate local and private serving of the open-weight release.
Limitations
- Text only: it does not natively process images, audio, or video.
- Non-commercial open-weight license: the CC-BY-NC-4.0 terms limit some commercial self-hosting and product uses.
- Significant hosting requirements: the 32-billion-parameter size can require substantial GPU memory without quantization.
- Not a specialist reasoning or coding model: users seeking advanced reasoning, software engineering, or agentic planning may need a more specialized option.
- Unspecified knowledge cutoff: official materials supplied for this review do not identify a cutoff date.
- Unverified tool features: function calling, streaming, batch processing, caching, and JSON mode are not confirmed by the supplied information.
When to choose Aya Expanse 32B
Choose Aya Expanse 32B when the central requirement is multilingual text quality across its supported languages. It is a sensible candidate for translation assistance, multilingual customer-service drafts, long-document summarization, international content production, education, and research. It is particularly attractive when a 128K context window and the possibility of non-commercial self-hosting are more important than maximum speed or specialized coding performance.
The API is generally the simpler choice when a team wants to test the model or integrate it without purchasing and operating GPU infrastructure. The open-weight route is more appropriate for eligible research and non-commercial deployments that require local control, customization, or experimentation with serving systems such as Transformers, vLLM, or SGLang.
Another type of model may be more appropriate when the application needs image understanding, audio transcription, video analysis, current web information, advanced coding, or verified tool and function support. A smaller model may also be preferable for high-volume, latency-sensitive workloads where multilingual breadth and a very large context window are less important. Conversely, teams requiring commercial self-hosting should review the license carefully and consider an option with terms that explicitly permit their planned deployment.
Bottom line
Aya Expanse 32B is best understood as a multilingual, text-only language model with a large context window and two deployment paths: Cohere’s token-priced API and a gated open-weight release for research and non-commercial use. Its value comes from combining 23-language coverage with long-context text generation, translation, and summarization. Its main trade-offs are substantial local infrastructure needs, a restrictive open-weight license for commercial scenarios, and a less specialized profile for multimodal work, advanced reasoning, and coding.

