Aya Expanse

Aya Expanse 32B

by Cohere · Live

Aya Expanse 32B is Cohere Labs’ 32-billion-parameter multilingual language model for text generation, translation, summarization, customer support, and long-document workflows. It supports 23 languages, offers a 128K-token context window and 4K maximum API output, and is available through Cohere’s API or as gated open weights for research and non-commercial use. The model is text-only, with significant local hosting requirements and no confirmed tool, JSON-mode, or advanced coding features.

Text Reasoning Coding
Cohere Labs released Aya Expanse 32B on October 24, 2024 as an open-weight multilingual language model. It combines instruction tuning, multilingual preference training, safety tuning, data-selection methods, and model merging to support consistent text generation across 23 languages. The model is available through Cohere’s Chat API and as gated open weights for research and non-commercial deployment.
Outputs

What Aya Expanse 32B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

6/10 Reasoning
5/10 Coding
4/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Aya Expanse
Model type General Purpose
Context window 128K tokens
Maximum output 4K tokens
Release date 2024-10-24
Status Live
Knowledge cutoff notes

The official Cohere model page and Hugging Face model card do not specify a knowledge-cutoff date for Aya Expanse 32B.

Model notes

Canonical Cohere API identifier is c4ai-aya-expanse-32b. The open-weight repository is CohereLabs/aya-expanse-32b and is gated. The model is text-only, supports 23 languages, has 32 billion parameters, and uses a 128K context window. Cohere documents a 4K maximum output for the Chat API. The open-weight release is licensed under CC-BY-NC-4.0 and also requires compliance with Cohere Labs’ acceptable-use policy. Fine-tuning is demonstrated through an official notebook, but hosted fine-tuning availability is not established. Editorial scores are comparative estimates, not provider-published benchmarks.

Cost

Model pricing

Input $0.50 per 1 million tokens
Output $1.50 per 1 million tokens
Model guide

Aya Expanse 32B: Cohere’s Open-Weight Multilingual Language Model

Aya Expanse 32B is Cohere Labs’ 32-billion-parameter, text-only language model built for multilingual generation, translation, summarization, customer support, and other language-intensive applications across 23 supported languages.

What is Aya Expanse 32B?

Aya Expanse 32B is a multilingual large language model from Cohere Labs, the research organization associated with Cohere. The model has 32 billion parameters and is designed primarily for text generation rather than image, audio, or video processing. Its canonical Cohere API identifier is c4ai-aya-expanse-32b, and its open-weight release is available through the gated Hugging Face repository CohereLabs/aya-expanse-32b.

The model is aimed at applications where language coverage matters as much as general text quality. Typical uses include translation assistance, multilingual writing, summarization, cross-language customer support, international communications, educational tools, and research involving multilingual language models. It can also be used for long-document work because its documented context window is 128,000 tokens.

Aya Expanse 32B fits within Cohere’s broader model lineup as a multilingual generation model. It is distinct from Cohere’s retrieval-focused Embed and Rerank families, document-processing tools such as Parse, and the company’s enterprise workspace products. Its primary role is to generate and transform text across many languages.

Languages and multilingual design

Aya Expanse 32B supports 23 languages: Arabic; simplified and traditional Chinese; Czech; Dutch; English; French; German; Greek; Hebrew; Hindi; Indonesian; Italian; Japanese; Korean; Persian; Polish; Portuguese; Romanian; Russian; Spanish; Turkish; Ukrainian; and Vietnamese.

Its main distinction is this broad multilingual focus, not multimodal input or specialized reasoning. Cohere Labs describes a training approach that combines supervised fine-tuning, multilingual preference training, safety tuning, data-arbitrage and data-selection techniques, and model merging. In practical terms, this is intended to help the model follow instructions and produce useful text across languages instead of concentrating most of its quality on English.

For example, a support application could use Aya Expanse 32B to summarize a customer request in one language and draft a response in another. A research workflow could use it to translate source material, create summaries, or compare information across languages. These are suitable applications because they use the model’s central strength: language transformation across a broad set of supported languages.

Context window and output limits

The model has a documented context window of 128,000 tokens. A context window is the amount of input and conversation history the model can consider in one request, measured in tokens rather than words. The large limit makes Aya Expanse 32B suitable for long documents, extended conversations, and multilingual material that would exceed the limits of smaller-context models.

Cohere documents a maximum API output of 4,000 tokens for the Chat API. The context limit and output limit serve different purposes: the 128K window covers the material supplied to the model, while the 4K maximum controls how much text it can generate in the response. The supplied documentation does not specify a knowledge-cutoff date for Aya Expanse 32B, so users should not assume that it contains current factual information.

API and open-weight deployment

There are two main ways to use Aya Expanse 32B. Developers can call it through Cohere’s Chat API using the model identifier c4ai-aya-expanse-32b. This avoids managing model files and hardware, but usage is billed according to token consumption and depends on the API arrangement and service availability.

The second option is the gated open-weight release on Hugging Face. The repository is intended for research and non-commercial use under the CC-BY-NC-4.0 license, together with Cohere Labs’ acceptable-use requirements. Access, licensing, and the intended deployment scenario should be reviewed before downloading or serving the model.

The open-weight model can be loaded with the Transformers ecosystem and served with tools such as vLLM or SGLang. A full-precision 32-billion-parameter model requires substantial GPU memory and inference capacity. Quantized community versions can reduce memory requirements, although quantization may reduce output quality. The supplied research does not establish a single hardware requirement, so actual needs will depend on precision, serving configuration, concurrency, and the chosen implementation.

API pricing

Cohere’s current pricing information lists Aya Expanse models, including Aya Expanse 32B, at $0.50 per million input tokens and $1.50 per million output tokens. Input tokens are the text sent to the model, while output tokens are the text it generates.

This is hosted API pricing, not a price for downloading or operating the open-weight release. Self-hosting introduces infrastructure, storage, engineering, and maintenance costs, even when the model weights themselves are available under the non-commercial license. Commercial users should verify whether the Cohere API or another authorized hosting arrangement permits their intended use; the open-weight CC-BY-NC-4.0 terms do not generally provide unrestricted commercial self-hosting rights.

Capabilities and supported modalities

Aya Expanse 32B is a text-in, text-out model. It accepts text prompts and generates text responses. The research identifies no native image, audio, or video input, and no image, audio, video, music, speech, or embedding output. It should therefore not be selected as a direct multimodal model for interpreting photographs, transcribing recordings, analyzing video, or generating media.

Its strongest supported tasks are multilingual text generation, translation, summarization, customer-support drafting, content creation, and global communication. The model can be incorporated into larger systems that provide document retrieval or other preprocessing, but those surrounding functions should not be confused with capabilities native to Aya Expanse 32B itself.

Tool use, function calling, streaming, caching, batch API support, and a distinct JSON mode are not established by the supplied research. Developers should not assume that these features are available simply because the model can be accessed through a Chat API. Structured application output may require application-side validation or a separately documented provider feature.

Reasoning, coding, speed, and cost trade-offs

Aya Expanse 32B is a general-purpose multilingual generator rather than a model specifically optimized for advanced reasoning or software engineering. It can produce explanations, transform text, and assist with ordinary language-oriented coding prompts, but the supplied research does not establish specialized coding performance or advanced reasoning benchmarks.

The model’s editorial reasoning score is 6 out of 10 and coding score is 5 out of 10. These are comparative editorial estimates, not provider-published benchmark results. They should be treated as directional judgments: Aya Expanse 32B is more compelling for multilingual language work than for demanding mathematical reasoning, complex agent planning, or code-heavy workflows.

Its editorial speed score is 4 out of 10 and cost score is 7 out of 10. These scores are also comparative estimates rather than verified provider metrics. The practical trade-off is straightforward: a 32-billion-parameter model with a 128K context window can provide substantial capacity for multilingual and long-context tasks, but it may require more compute and respond more slowly than smaller models. API users pay relatively low token rates in the supplied pricing information, while self-hosting users must account for GPU and operational costs.

Main strengths and limitations

Strengths

  • Broad language coverage: it supports 23 languages, making it useful for cross-language workflows and international applications.
  • Long context: the 128K context window supports large documents and extended interactions.
  • Open-weight availability: researchers and eligible non-commercial users can examine and deploy the gated release rather than relying exclusively on a hosted endpoint.
  • Useful text tasks: translation, summarization, multilingual writing, customer support, and international communication align closely with its intended role.
  • Flexible deployment choices: users can select Cohere’s API or investigate local and private serving of the open-weight release.

Limitations

  • Text only: it does not natively process images, audio, or video.
  • Non-commercial open-weight license: the CC-BY-NC-4.0 terms limit some commercial self-hosting and product uses.
  • Significant hosting requirements: the 32-billion-parameter size can require substantial GPU memory without quantization.
  • Not a specialist reasoning or coding model: users seeking advanced reasoning, software engineering, or agentic planning may need a more specialized option.
  • Unspecified knowledge cutoff: official materials supplied for this review do not identify a cutoff date.
  • Unverified tool features: function calling, streaming, batch processing, caching, and JSON mode are not confirmed by the supplied information.

When to choose Aya Expanse 32B

Choose Aya Expanse 32B when the central requirement is multilingual text quality across its supported languages. It is a sensible candidate for translation assistance, multilingual customer-service drafts, long-document summarization, international content production, education, and research. It is particularly attractive when a 128K context window and the possibility of non-commercial self-hosting are more important than maximum speed or specialized coding performance.

The API is generally the simpler choice when a team wants to test the model or integrate it without purchasing and operating GPU infrastructure. The open-weight route is more appropriate for eligible research and non-commercial deployments that require local control, customization, or experimentation with serving systems such as Transformers, vLLM, or SGLang.

Another type of model may be more appropriate when the application needs image understanding, audio transcription, video analysis, current web information, advanced coding, or verified tool and function support. A smaller model may also be preferable for high-volume, latency-sensitive workloads where multilingual breadth and a very large context window are less important. Conversely, teams requiring commercial self-hosting should review the license carefully and consider an option with terms that explicitly permit their planned deployment.

Bottom line

Aya Expanse 32B is best understood as a multilingual, text-only language model with a large context window and two deployment paths: Cohere’s token-priced API and a gated open-weight release for research and non-commercial use. Its value comes from combining 23-language coverage with long-context text generation, translation, and summarization. Its main trade-offs are substantial local infrastructure needs, a restrictive open-weight license for commercial scenarios, and a less specialized profile for multimodal work, advanced reasoning, and coding.


Answers to Frequently Asked Questions

How much does Aya Expanse 32B cost, and can it be used commercially?
Cohere’s listed API pricing is $0.50 per million input tokens and $1.50 per million output tokens. The open-weight release is available under the CC-BY-NC-4.0 license for research and non-commercial use, so commercial self-hosting and product use require careful license review.
How can I access and deploy Aya Expanse 32B?
Developers can access Aya Expanse 32B through Cohere’s Chat API using the identifier c4ai-aya-expanse-32b, or use the gated open-weight release at CohereLabs/aya-expanse-32b on Hugging Face. The open-weight model can be loaded with Transformers and served with tools such as vLLM or SGLang, but it requires substantial GPU capacity.
What is the context window and output limit of Aya Expanse 32B?
The model has a documented context window of 128,000 tokens, making it suitable for long documents and extended conversations. Cohere documents a maximum API output of 4,000 tokens for the Chat API.
What is Aya Expanse 32B?
Aya Expanse 32B is a 32-billion-parameter multilingual, text-only language model from Cohere Labs. It is designed for text generation, translation, summarization, multilingual writing, and cross-language communication.
Which languages does Aya Expanse 32B support?
Aya Expanse 32B supports 23 languages: Arabic, simplified and traditional Chinese, Czech, Dutch, English, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Korean, Persian, Polish, Portuguese, Romanian, Russian, Spanish, Turkish, Ukrainian, and Vietnamese.


Sources 6
Provider

About Cohere