What is Qwen3-235B-A22B?
Qwen3-235B-A22B is a large language model from Alibaba Cloud’s Qwen family. It is built for text-based tasks such as reasoning, software development, multilingual writing, research assistance and agentic applications. The model was released on April 29, 2025, and is available both as an open-weight model and through Alibaba Cloud Model Studio.
The name describes its scale: the model has approximately 235 billion total parameters, while about 22 billion parameters are activated for each token generated. This is a mixture-of-experts (MoE) design. Instead of running every part of the network for every token, a routing system selects a subset of specialized experts. That reduces the computation required for each step compared with a dense model of the same total size, although the model still has substantial infrastructure requirements when self-hosted.
Qwen3-235B-A22B is the original model rather than one of the later separately identified Qwen3-235B-A22B-Instruct-2507 or Qwen3-235B-A22B-Thinking-2507 variants. This distinction matters when comparing model cards, pricing and behavior.
Architecture, context and output limits
| Specification | Verified detail |
|---|---|
| Provider | Alibaba Cloud |
| Model family | Qwen3 |
| Architecture | Sparse mixture of experts |
| Total parameters | 235 billion |
| Active parameters | Approximately 22 billion per token |
| Routed experts | 128 |
| Activated experts | 8 per token |
| Context window | 131,072 tokens |
| Maximum output | 16,384 tokens |
| Open-weight license | Apache 2.0 |
The 131,072-token context window is useful for long documents, large codebases, extended conversations and multi-step research prompts. It is a maximum context limit, not a guarantee that every request will receive equally strong results at that length. Applications should still manage retrieval, prompt structure and output size carefully.
The maximum output limit is 16,384 tokens. That gives the model room for detailed explanations, code generation and reasoning traces where enabled, but it does not mean every response will use the full limit. Actual limits can also depend on the service surface and request configuration.
Thinking and non-thinking modes
A defining feature of Qwen3-235B-A22B is its hybrid operation. The same original model identity can be used in a reasoning-oriented thinking mode or a faster non-thinking mode. Thinking mode is intended for tasks that benefit from additional intermediate reasoning, such as difficult mathematics, multi-step analysis, debugging and planning. Non-thinking mode is better suited to straightforward questions, routine transformations and applications where response latency matters more.
This creates a practical control rather than requiring a separate model for every workload. A developer can reserve thinking mode for difficult requests and use non-thinking mode for simpler turns. The trade-off is that thinking-mode output is priced higher in Alibaba Cloud Model Studio, and responses may take longer than standard generation.
The supplied reasoning and coding ratings of 9 out of 10 are editorial comparative estimates, not Alibaba Cloud benchmark scores or official specifications. They indicate the model’s intended positioning and observed capability assessment in the supplied research, but they should not be treated as provider-published performance guarantees.
Capabilities and supported inputs
Qwen3-235B-A22B is a text-only model in the documented Model Studio configuration. It accepts text input and produces text output. It does not provide image, audio or video understanding or generation. The broader Qwen product ecosystem advertises multimodal products, but those provider-level features should not be attributed to this specific model.
Its supported uses include:
- Reasoning: multi-step analysis, mathematics, planning and difficult question answering, particularly in thinking mode.
- Coding: code generation, explanation, refactoring, debugging and software-development assistance.
- Multilingual generation: writing and transformation across supported languages, including use cases where multilingual capability is important.
- Structured generation: structured outputs for applications that need responses in a specified format.
- Function calling: selecting and invoking application-defined tools as part of an agent workflow.
- Long-context work: analysis of large documents, specifications, transcripts or code-related material within the context limit.
Function calling does not mean the model independently has access to the internet, databases or private systems. The application must define the available tools, execute calls and return results to the model. The Model Studio documentation identifies web search as unsupported for this model, so an application needing live information must provide an appropriate external tool rather than assuming built-in browsing.
Pricing through Alibaba Cloud Model Studio
Alibaba Cloud’s documented pricing varies by region, mode and whether tokens are input or output. For the United States, Germany and China, standard input is listed at $0.287 per 1 million tokens, while standard output is $1.147 per 1 million tokens. Thinking-mode input is priced at the same input rate, but thinking-mode output is listed at $2.868 per 1 million tokens.
Singapore pricing is higher: standard input is $0.700 per 1 million tokens, standard output is $2.800 per 1 million tokens, and thinking-mode output is $8.400 per 1 million tokens. These are usage prices rather than a consumer subscription fee, and the applicable region and service terms should be checked before deployment.
| Region | Standard input | Standard output | Thinking output |
|---|---|---|---|
| United States, Germany and China | $0.287 per 1M tokens | $1.147 per 1M tokens | $2.868 per 1M tokens |
| Singapore | $0.700 per 1M tokens | $2.800 per 1M tokens | $8.400 per 1M tokens |
Because output is more expensive than input and thinking output costs more than standard output, cost control should focus on routing simple requests to non-thinking mode, limiting unnecessary context and setting sensible output caps. The supplied editorial cost score is 8 out of 10; this is a comparative estimate, not a vendor rating.
Tools, development and deployment
At the hosted API level, Qwen3-235B-A22B supports function calling, streaming and structured outputs. Streaming is useful for interactive interfaces because the application can display generated text incrementally rather than waiting for the complete response. Structured outputs can make it easier to connect the model to downstream software, although the supplied research does not verify a separate product feature specifically named “JSON mode.”
The open-weight release can be self-hosted or adapted with compatible tools such as Transformers, vLLM and SGLang. The Apache 2.0 license supports a broad range of deployment and modification scenarios, subject to the license and any applicable operational or legal requirements. The model is therefore more flexible than a hosted-only service, but running a model of this scale requires substantially more engineering and hardware planning than using a smaller model through an API.
Alibaba Cloud Model Studio documents unsupported context caching, batch inference and hosted fine-tuning for this model. That does not prevent users of the open-weight release from fine-tuning or adapting it with compatible infrastructure; it means those capabilities should not be assumed to be managed features of the documented hosted offering.
Main strengths and limitations
Where the model is strong
- Its MoE architecture combines very high total parameter capacity with a lower active-parameter count per token.
- Thinking and non-thinking modes allow a practical balance between deeper reasoning and faster responses.
- The 131K context window supports substantial documents, code and multi-step prompts.
- Text reasoning, coding, multilingual generation, structured outputs and function calling cover many demanding application workflows.
- The open-weight Apache 2.0 release supports self-hosting, customization and deployment outside a single managed API.
Where it is less suitable
- It is text-only and should not be selected for native image, audio or video input or output.
- It is not positioned as an ultra-low-latency model; the supplied editorial speed score is 6 out of 10.
- Thinking-mode output is considerably more expensive than standard output, especially in Singapore.
- Model Studio does not document built-in web search, context caching, batch inference or hosted fine-tuning for this model.
- The 235B total-parameter scale makes self-hosting more demanding than deploying a smaller model.
- There is no authoritative knowledge-cutoff date identified in the supplied research, so current facts should be supplied through verified tools or sources rather than assumed to be in the model’s training data.
When to choose Qwen3-235B-A22B
Choose Qwen3-235B-A22B when the task benefits from strong reasoning, coding ability, a long context window and the option to deploy open weights. It is a good candidate for complex software-development assistance, mathematical or analytical workflows, multilingual applications, document-heavy research systems and agents that need function calling.
Its hybrid modes are especially useful when one application serves different request types. Routine classification, rewriting or short answers can use non-thinking mode, while difficult planning or debugging requests can use thinking mode. This can provide better cost and latency control than sending every request through the most expensive reasoning setting.
A smaller model may be more appropriate for high-volume, latency-sensitive or resource-constrained workloads. A multimodal model is a better choice when the application must inspect images, audio or video. A hosted model with managed batch processing, prompt caching or web search may also be preferable when those service features are more important than open-weight deployment. Finally, one of the later Qwen3-235B-A22B 2507 variants may be worth evaluating when the required behavior specifically matches its separate instruct or thinking release, but those variants should not be treated as interchangeable with the original model.
Bottom line
Qwen3-235B-A22B is aimed at demanding text workloads where reasoning depth, coding, long context and deployment flexibility matter more than minimum latency. Its main practical distinction is the combination of a 235B-total-parameter MoE design, approximately 22B active parameters per token, switchable thinking behavior and an open-weight Apache 2.0 release. The main compromises are text-only operation, substantial self-hosting requirements, higher thinking-mode output costs and the absence of several managed API features documented for other service configurations.

