Qwen3

Qwen3-235B-A22B

by Qwen · Current and accessible through Alibaba Cloud Model Studio; original Qwen3 open-weight release, with newer 2507 instruct and thinking variants available separately

Qwen3-235B-A22B is Alibaba Cloud’s original Qwen3 flagship: an Apache 2.0 open-weight mixture-of-experts model with 235B total and approximately 22B active parameters. It provides thinking and non-thinking modes, a 131K-token context window, 16,384-token maximum output, text-only processing, coding and reasoning capabilities, function calling, structured outputs and regional Model Studio pricing.

Text Reasoning Coding
Released on April 29, 2025, Qwen3-235B-A22B is the original flagship model in the Qwen3 family. Its sparse mixture-of-experts architecture provides the capacity of a 235-billion-parameter model while activating about 22 billion parameters for each token. The model is available as an Apache 2.0 open-weight release and through Alibaba Cloud Model Studio, where it supports text generation, hybrid thinking modes, structured outputs and tool or function calling.
Outputs

What Qwen3-235B-A22B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning Structured output
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3
Model type Reasoning
Context window 131K tokens
Maximum output 16K tokens
Release date 2025-04-29
Status Current and accessible through Alibaba Cloud Model Studio; original Qwen3 open-weight release, with newer 2507 instruct and thinking variants available separately
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was identified for the exact original Qwen3-235B-A22B model. Release date, context length and API limits should not be interpreted as a knowledge cutoff.

Model notes

Qwen3-235B-A22B is a sparse MoE model with 235B total parameters, 22B active parameters, 128 routed experts and 8 activated experts per token. The original model supports hybrid thinking and non-thinking modes through the same model identity. Alibaba Cloud Model Studio documents text-only input and output, function calling, structured outputs and unsupported web search, context caching, batch inference and hosted fine-tuning. The open-weight release is available under Apache 2.0 and can be fine-tuned or self-hosted using compatible tooling such as Transformers, vLLM, SGLang and related Qwen ecosystem tools. The original open-weight model is distinct from the later Qwen3-235B-A22B-Instruct-2507 and Qwen3-235B-A22B-Thinking-2507 variants. Reasoning, coding, speed and cost scores are editorial comparative estimates rather than vendor specifications.

Cost

Model pricing

Input $0.287 per 1M tokens for standard input in the United States, Germany and China; $0.700 per 1M tokens in Singapore. Thinking-mode input is priced the same.
Output $1.147 per 1M tokens for standard output in the United States, Germany and China; $2.868 per 1M tokens for thinking-mode output. Singapore pricing is $2.800 standard output and $8.400 thinking-mode output per 1M tokens.
Model guide

Qwen3-235B-A22B: Open-Weight MoE Reasoning for Coding and Agents

Qwen3-235B-A22B is Alibaba Cloud’s open-weight mixture-of-experts language model with 235 billion total parameters and approximately 22 billion active parameters per token. It combines a reasoning-focused thinking mode with a faster non-thinking mode, supports a 131,072-token context window and 16,384-token maximum output, and is designed for complex reasoning, coding, multilingual generation, function calling and agentic workflows.

What is Qwen3-235B-A22B?

Qwen3-235B-A22B is a large language model from Alibaba Cloud’s Qwen family. It is built for text-based tasks such as reasoning, software development, multilingual writing, research assistance and agentic applications. The model was released on April 29, 2025, and is available both as an open-weight model and through Alibaba Cloud Model Studio.

The name describes its scale: the model has approximately 235 billion total parameters, while about 22 billion parameters are activated for each token generated. This is a mixture-of-experts (MoE) design. Instead of running every part of the network for every token, a routing system selects a subset of specialized experts. That reduces the computation required for each step compared with a dense model of the same total size, although the model still has substantial infrastructure requirements when self-hosted.

Qwen3-235B-A22B is the original model rather than one of the later separately identified Qwen3-235B-A22B-Instruct-2507 or Qwen3-235B-A22B-Thinking-2507 variants. This distinction matters when comparing model cards, pricing and behavior.

Architecture, context and output limits

SpecificationVerified detail
ProviderAlibaba Cloud
Model familyQwen3
ArchitectureSparse mixture of experts
Total parameters235 billion
Active parametersApproximately 22 billion per token
Routed experts128
Activated experts8 per token
Context window131,072 tokens
Maximum output16,384 tokens
Open-weight licenseApache 2.0

The 131,072-token context window is useful for long documents, large codebases, extended conversations and multi-step research prompts. It is a maximum context limit, not a guarantee that every request will receive equally strong results at that length. Applications should still manage retrieval, prompt structure and output size carefully.

The maximum output limit is 16,384 tokens. That gives the model room for detailed explanations, code generation and reasoning traces where enabled, but it does not mean every response will use the full limit. Actual limits can also depend on the service surface and request configuration.

Thinking and non-thinking modes

A defining feature of Qwen3-235B-A22B is its hybrid operation. The same original model identity can be used in a reasoning-oriented thinking mode or a faster non-thinking mode. Thinking mode is intended for tasks that benefit from additional intermediate reasoning, such as difficult mathematics, multi-step analysis, debugging and planning. Non-thinking mode is better suited to straightforward questions, routine transformations and applications where response latency matters more.

This creates a practical control rather than requiring a separate model for every workload. A developer can reserve thinking mode for difficult requests and use non-thinking mode for simpler turns. The trade-off is that thinking-mode output is priced higher in Alibaba Cloud Model Studio, and responses may take longer than standard generation.

The supplied reasoning and coding ratings of 9 out of 10 are editorial comparative estimates, not Alibaba Cloud benchmark scores or official specifications. They indicate the model’s intended positioning and observed capability assessment in the supplied research, but they should not be treated as provider-published performance guarantees.

Capabilities and supported inputs

Qwen3-235B-A22B is a text-only model in the documented Model Studio configuration. It accepts text input and produces text output. It does not provide image, audio or video understanding or generation. The broader Qwen product ecosystem advertises multimodal products, but those provider-level features should not be attributed to this specific model.

Its supported uses include:

  • Reasoning: multi-step analysis, mathematics, planning and difficult question answering, particularly in thinking mode.
  • Coding: code generation, explanation, refactoring, debugging and software-development assistance.
  • Multilingual generation: writing and transformation across supported languages, including use cases where multilingual capability is important.
  • Structured generation: structured outputs for applications that need responses in a specified format.
  • Function calling: selecting and invoking application-defined tools as part of an agent workflow.
  • Long-context work: analysis of large documents, specifications, transcripts or code-related material within the context limit.

Function calling does not mean the model independently has access to the internet, databases or private systems. The application must define the available tools, execute calls and return results to the model. The Model Studio documentation identifies web search as unsupported for this model, so an application needing live information must provide an appropriate external tool rather than assuming built-in browsing.

Pricing through Alibaba Cloud Model Studio

Alibaba Cloud’s documented pricing varies by region, mode and whether tokens are input or output. For the United States, Germany and China, standard input is listed at $0.287 per 1 million tokens, while standard output is $1.147 per 1 million tokens. Thinking-mode input is priced at the same input rate, but thinking-mode output is listed at $2.868 per 1 million tokens.

Singapore pricing is higher: standard input is $0.700 per 1 million tokens, standard output is $2.800 per 1 million tokens, and thinking-mode output is $8.400 per 1 million tokens. These are usage prices rather than a consumer subscription fee, and the applicable region and service terms should be checked before deployment.

RegionStandard inputStandard outputThinking output
United States, Germany and China$0.287 per 1M tokens$1.147 per 1M tokens$2.868 per 1M tokens
Singapore$0.700 per 1M tokens$2.800 per 1M tokens$8.400 per 1M tokens

Because output is more expensive than input and thinking output costs more than standard output, cost control should focus on routing simple requests to non-thinking mode, limiting unnecessary context and setting sensible output caps. The supplied editorial cost score is 8 out of 10; this is a comparative estimate, not a vendor rating.

Tools, development and deployment

At the hosted API level, Qwen3-235B-A22B supports function calling, streaming and structured outputs. Streaming is useful for interactive interfaces because the application can display generated text incrementally rather than waiting for the complete response. Structured outputs can make it easier to connect the model to downstream software, although the supplied research does not verify a separate product feature specifically named “JSON mode.”

The open-weight release can be self-hosted or adapted with compatible tools such as Transformers, vLLM and SGLang. The Apache 2.0 license supports a broad range of deployment and modification scenarios, subject to the license and any applicable operational or legal requirements. The model is therefore more flexible than a hosted-only service, but running a model of this scale requires substantially more engineering and hardware planning than using a smaller model through an API.

Alibaba Cloud Model Studio documents unsupported context caching, batch inference and hosted fine-tuning for this model. That does not prevent users of the open-weight release from fine-tuning or adapting it with compatible infrastructure; it means those capabilities should not be assumed to be managed features of the documented hosted offering.

Main strengths and limitations

Where the model is strong

  • Its MoE architecture combines very high total parameter capacity with a lower active-parameter count per token.
  • Thinking and non-thinking modes allow a practical balance between deeper reasoning and faster responses.
  • The 131K context window supports substantial documents, code and multi-step prompts.
  • Text reasoning, coding, multilingual generation, structured outputs and function calling cover many demanding application workflows.
  • The open-weight Apache 2.0 release supports self-hosting, customization and deployment outside a single managed API.

Where it is less suitable

  • It is text-only and should not be selected for native image, audio or video input or output.
  • It is not positioned as an ultra-low-latency model; the supplied editorial speed score is 6 out of 10.
  • Thinking-mode output is considerably more expensive than standard output, especially in Singapore.
  • Model Studio does not document built-in web search, context caching, batch inference or hosted fine-tuning for this model.
  • The 235B total-parameter scale makes self-hosting more demanding than deploying a smaller model.
  • There is no authoritative knowledge-cutoff date identified in the supplied research, so current facts should be supplied through verified tools or sources rather than assumed to be in the model’s training data.

When to choose Qwen3-235B-A22B

Choose Qwen3-235B-A22B when the task benefits from strong reasoning, coding ability, a long context window and the option to deploy open weights. It is a good candidate for complex software-development assistance, mathematical or analytical workflows, multilingual applications, document-heavy research systems and agents that need function calling.

Its hybrid modes are especially useful when one application serves different request types. Routine classification, rewriting or short answers can use non-thinking mode, while difficult planning or debugging requests can use thinking mode. This can provide better cost and latency control than sending every request through the most expensive reasoning setting.

A smaller model may be more appropriate for high-volume, latency-sensitive or resource-constrained workloads. A multimodal model is a better choice when the application must inspect images, audio or video. A hosted model with managed batch processing, prompt caching or web search may also be preferable when those service features are more important than open-weight deployment. Finally, one of the later Qwen3-235B-A22B 2507 variants may be worth evaluating when the required behavior specifically matches its separate instruct or thinking release, but those variants should not be treated as interchangeable with the original model.

Bottom line

Qwen3-235B-A22B is aimed at demanding text workloads where reasoning depth, coding, long context and deployment flexibility matter more than minimum latency. Its main practical distinction is the combination of a 235B-total-parameter MoE design, approximately 22B active parameters per token, switchable thinking behavior and an open-weight Apache 2.0 release. The main compromises are text-only operation, substantial self-hosting requirements, higher thinking-mode output costs and the absence of several managed API features documented for other service configurations.


Answers to Frequently Asked Questions

Can Qwen3-235B-A22B be self-hosted?
Yes. Qwen3-235B-A22B is released under the Apache 2.0 license and can be self-hosted or adapted with compatible tools such as Transformers, vLLM and SGLang. However, its 235-billion-parameter scale requires substantial hardware and engineering resources compared with smaller models.
Does Qwen3-235B-A22B support multimodal input or web search?
No. In the documented Model Studio configuration, Qwen3-235B-A22B is text-only and does not natively process images, audio or video. It also does not include built-in web search. Applications that need live information must connect external tools through function calling.
What is the difference between thinking and non-thinking modes in Qwen3-235B-A22B?
Thinking mode is intended for complex reasoning, mathematics, debugging and planning, while non-thinking mode is better for routine tasks that require lower latency. Thinking-mode output costs more and may take longer, so applications can route difficult requests to thinking mode and simpler requests to non-thinking mode.
What are the context window and output limits of Qwen3-235B-A22B?
Qwen3-235B-A22B supports a context window of up to 131,072 tokens and a maximum output of 16,384 tokens. These limits support long documents, large codebases and multi-step prompts, although actual service limits may depend on the deployment and request configuration.
What is Qwen3-235B-A22B?
Qwen3-235B-A22B is an open-weight large language model from Alibaba Cloud’s Qwen3 family, designed for text-based reasoning, coding, multilingual generation, research assistance and agentic applications. It uses a sparse mixture-of-experts architecture with approximately 235 billion total parameters and about 22 billion active parameters per token.


Sources 6
Provider

About Qwen