Qwen3.8

Qwen3.8-2.4T-A95B

by Qwen · Current and available; open-weight release and hosted API model

Alibaba’s Qwen3.8-2.4T-A95B is an open-weight flagship mixture-of-experts language model for advanced reasoning, coding, research, and long-horizon agents. It offers a 1-million-token context window, 131,072-token maximum output, thinking and non-thinking modes, function calling, structured outputs, web search, streaming, and context caching. The text-only model is expensive and infrastructure-intensive, with hosted batch inference and fine-tuning unsupported for this exact release.

Text Reasoning Coding
Qwen3.8-2.4T-A95B is Alibaba’s open-weight flagship model for demanding reasoning, coding, research, professional-work, and long-horizon agent tasks. It combines a sparse mixture-of-experts architecture with a documented 1-million-token context window and supports both downloadable weights and a hosted Alibaba Cloud Model Studio endpoint. The model is text-only, and its very large parameter count means that organizations must weigh its capability against infrastructure, latency, and operating-cost requirements.
Outputs

What Qwen3.8-2.4T-A95B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching
Model profile

Performance characteristics

10/10 Reasoning
10/10 Coding
4/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.8
Model type Reasoning
Context window 1M tokens
Maximum output 131K tokens
Release date 2026-08-12
Status Current and available; open-weight release and hosted API model
Knowledge cutoff notes

No authoritative first-party knowledge-cutoff date was identified for this exact model.

Model notes

Qwen3.8-2.4T-A95B is the open-weight Qwen3.8 flagship model with approximately 2.4 trillion total parameters and about 95 billion activated parameters per step. The official hosted API identifier is qwen3.8-2.4t-a95b, while the downloadable model identifier is Qwen/Qwen3.8-2.4T-A95B. Model Studio supports thinking and non-thinking modes, function calling, structured outputs, web search, context caching, and streaming. Model Studio documents batch inference and fine-tuning as unsupported for this exact model. International cache-hit pricing is $0.25 per 1 million tokens; explicit cache creation is $2.50 per 1 million tokens and explicit cache hits are $0.17 per 1 million tokens. The model is text-only at the model endpoint despite broader multimodal capabilities elsewhere in the Qwen ecosystem. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input $2 per 1 million tokens internationally; $1.65 per 1 million tokens in China (Beijing)
Output $6 per 1 million tokens internationally; $4.951 per 1 million tokens in China (Beijing)
Model guide

Qwen3.8-2.4T-A95B: Alibaba’s Open-Weight Model for Long-Horizon Reasoning

Qwen3.8-2.4T-A95B is Alibaba’s open-weight flagship language model, built as a sparse mixture-of-experts system with approximately 2.4 trillion total parameters and about 95 billion active parameters per step. It supports a 1-million-token context window, reasoning and non-thinking modes, tool use, structured outputs, web search, streaming, and context caching. Its capability and long-context focus make it suitable for advanced reasoning, coding, research, and agent workflows, while its scale makes self-hosting difficult and hosted use relatively expensive.

What is Qwen3.8-2.4T-A95B?

Qwen3.8-2.4T-A95B is an open-weight text-generation model from Alibaba’s Qwen family. The name reflects its approximate scale: the model contains 2.4 trillion total parameters, although its sparse mixture-of-experts (MoE) design activates about 95 billion parameters for an individual step. In an MoE model, only selected expert networks process each token rather than the entire parameter set, allowing a very large model to use computation more selectively.

Alibaba positions this model for advanced reasoning, software engineering, scientific and professional research, large-document analysis, and long-horizon agent workflows. It is available as downloadable Qwen weights and through Alibaba Cloud Model Studio, where the hosted API identifier is qwen3.8-2.4t-a95b. The downloadable checkpoint is identified as Qwen/Qwen3.8-2.4T-A95B.

This is a model-level product, not a general description of every capability in Qwen Studio. Features offered elsewhere in the Qwen ecosystem should not be assumed to be available through this specific endpoint.

Architecture and context limits

The model combines sparse MoE routing with hybrid attention. Its documented context length is 1,000,000 tokens, giving it an unusually large working window for long documents, codebases, research collections, and extended agent histories. The maximum output length is 131,072 tokens.

For reasoning use, Model Studio documents a maximum input length of 983,616 tokens and a maximum chain-of-thought length of 131,072 tokens in thinking mode. The exact usable limit can depend on the selected mode and the provider’s serving configuration, so applications should reserve room for the requested output rather than treating the full context figure as available input space.

A million-token context window does not automatically mean that every long prompt will produce equally reliable results. Very large prompts can increase processing time and cost, and applications still need to manage retrieval quality, prompt organization, and irrelevant material.

Reasoning, coding, and tool support

Qwen3.8-2.4T-A95B supports thinking and non-thinking inference modes. Thinking mode is intended for tasks that benefit from more deliberate intermediate reasoning, while non-thinking mode can be used when a shorter or more direct response is preferable. The supplied research supports its use for complex problem solving, coding, research, and professional knowledge work, but does not provide a provider-published benchmark score for this exact model.

For software work, the model supports text-based code generation and is positioned for advanced software engineering. Its large context window can be useful for examining extensive specifications, multiple source files, logs, or documentation in one request. As with other code-generating systems, generated code should be reviewed and tested rather than treated as automatically correct.

Model Studio supports function calling and tool use. It also documents structured outputs, allowing applications to request responses in a defined structure instead of relying only on free-form text parsing. Structured outputs should not be confused with a separately verified JSON-mode guarantee; the supplied model data does not specify a distinct JSON-mode capability.

Hosted Model Studio access also supports first-party web search, streaming, and context caching. Web search and other tools are service-level features around the model endpoint, not additional native output modalities. Context caching can reduce the cost of repeatedly processing shared prompt material, although cache-hit and cache-creation rates differ from standard input pricing.

Supported input and output modalities

The documented endpoint is text-only. It accepts text input and produces text output. It does not natively provide image, audio, or video input or output according to the supplied model specifications.

Qwen Studio and the broader Qwen ecosystem advertise multimodal features, including image, audio, and video understanding or generation in other products and models. Those ecosystem capabilities should not be attributed to Qwen3.8-2.4T-A95B itself. This distinction matters when selecting an endpoint: a workflow that needs image analysis, speech processing, or media generation requires a separately supported model or service.

Pricing and availability

Alibaba Cloud Model Studio lists international standard pricing of $2 per 1 million input tokens and $6 per 1 million output tokens for this model. Singapore pricing also includes separate cache-related rates. In China, the supplied pricing lists $1.65 per 1 million input tokens and $4.951 per 1 million output tokens in Beijing. Regional prices, eligibility, and service availability should be checked before deployment.

For international usage, the supplied research lists cache-hit pricing of $0.25 per 1 million tokens, explicit cache creation at $2.50 per 1 million tokens, and explicit cache hits at $0.17 per 1 million tokens. These rates are alternatives for specific caching operations, not additional charges to combine indiscriminately with the standard token rates.

Model Studio identifies batch inference and fine-tuning as unsupported for this exact model. The downloadable open-weight release provides a separate self-hosting path, but the model’s scale makes deployment substantially more demanding than smaller language models. Quantized checkpoints may reduce memory requirements, but the supplied research does not establish a specific hardware configuration or guarantee a particular performance level.

Main strengths and trade-offs

  • Very large context: The 1-million-token window is suited to long documents, extensive code repositories, and prolonged agent histories.
  • High-capability positioning: Alibaba targets the model at advanced reasoning, coding, research, and professional workloads.
  • Flexible reasoning: Thinking and non-thinking modes let applications trade response depth against directness and likely latency.
  • Developer features: Function calling, structured outputs, streaming, web search, and context caching support application workflows.
  • Open-weight availability: Organizations can evaluate downloadable weights rather than relying only on a hosted endpoint.
  • Infrastructure burden: The 2.4-trillion-parameter scale makes full self-hosting practical mainly for organizations with substantial accelerator infrastructure.
  • Cost and speed: The model is not designed as a low-cost, low-latency option. Its editorial speed score is comparatively low and its cost score is moderate; these are editorial evaluations, not Alibaba-published benchmarks.
  • Limited service options: Hosted batch inference and fine-tuning are not documented as supported for this exact model.
  • Text-only endpoint: Native image, audio, and video input or output are not supported.

Best use cases

Qwen3.8-2.4T-A95B is a strong candidate when the task benefits from deep reasoning, a large working context, or extended tool interaction. Suitable examples include:

  • Analyzing large technical, legal, scientific, or business document collections.
  • Reviewing large codebases and producing coordinated changes across many files.
  • Building research assistants that combine reasoning with web search.
  • Orchestrating long-running agents that need to retain substantial task history.
  • Generating or reviewing complex software and technical documentation.
  • Deploying an open-weight flagship model in an organization with the infrastructure to operate it.

For repeated prompts containing the same large reference material, context caching may improve the economics of hosted use. For interactive applications where response time is more important than maximum reasoning capacity, a smaller model may be more appropriate.

When to choose this model

Choose Qwen3.8-2.4T-A95B when maximum capability within the Qwen3.8 offering, long-context processing, and advanced reasoning matter more than minimal cost or latency. It is particularly relevant for organizations that can use multi-GPU infrastructure or that want to evaluate an open-weight flagship model while retaining the option of hosted access through Model Studio.

Consider another option when the workload needs fast conversational responses, consumer-hardware deployment, predictable low operating costs, native media understanding or generation, hosted fine-tuning, or batch processing. A smaller language model will generally be easier and cheaper to operate, while a dedicated multimodal model is a better fit for image, audio, or video workflows. The supplied research does not identify a specific sibling model with directly comparable pricing and limits, so those alternatives should be evaluated separately rather than inferred from this model’s specifications.

Bottom line

Qwen3.8-2.4T-A95B is aimed at demanding workloads rather than everyday lightweight chat. Its defining practical characteristics are the sparse 2.4-trillion-parameter design, approximately 95 billion active parameters per step, 1-million-token context window, reasoning modes, developer tooling, and open-weight availability. Those benefits come with substantial deployment requirements, relatively high hosted token prices, and the limitations of a text-only model without hosted fine-tuning or batch inference. It is best evaluated as a high-end reasoning and long-context system, not as a general-purpose replacement for every model in the Qwen ecosystem.


Answers to Frequently Asked Questions

How much does Qwen3.8-2.4T-A95B cost on Alibaba Cloud Model Studio?
International standard pricing is listed as $2 per 1 million input tokens and $6 per 1 million output tokens. Regional pricing and cache-related rates vary, so current Model Studio pricing and availability should be checked before deployment.
Does Qwen3.8-2.4T-A95B support multimodal input and output?
No. The documented endpoint is text-only, accepting text input and producing text output. It does not natively support image, audio, or video input or output, despite multimodal capabilities being available in other Qwen products and models.
What can Qwen3.8-2.4T-A95B be used for?
The model is designed for advanced reasoning, software engineering, scientific and professional research, large-document analysis, codebase review, and long-running agent workflows. Its large context window is especially useful for extensive documents, repositories, logs, and research collections.
What is Qwen3.8-2.4T-A95B?
Qwen3.8-2.4T-A95B is an open-weight text-generation model from Alibaba’s Qwen family. It uses a sparse mixture-of-experts architecture with approximately 2.4 trillion total parameters and about 95 billion active parameters per step.
How large is Qwen3.8-2.4T-A95B’s context window?
The model has a documented context length of 1,000,000 tokens and a maximum output length of 131,072 tokens. In thinking mode, Model Studio documents a maximum input length of 983,616 tokens and a maximum chain-of-thought length of 131,072 tokens.


Sources 7
Provider

About Qwen