Qwen3.8

Qwen3.8-Max

by Qwen · Stable official release; currently available through Alibaba Cloud Model Studio

Alibaba Cloud’s Qwen3.8-Max is a flagship multimodal mixture-of-experts model for large documents, long videos, complex coding, visual reasoning, and tool-assisted agents. It accepts text, images, and video, supports thinking mode, function calling, web search, structured output, streaming, and context caching, and provides a one-million-token context window with up to 131,072 output tokens. Its main trade-offs are higher cost, a less lightweight speed profile, no native media generation, and no fine-tuning support in the model-specific capability record.

Text Reasoning Coding
Qwen3.8-Max is Alibaba Cloud’s largest and most capable Qwen model in the supplied lineup. Released on August 3, 2026, it is designed for tasks that require more than a short answer: analyzing very large documents, understanding long videos, solving complex coding problems, using tools across multiple steps, and supporting professional research or software-engineering workflows. The model accepts text, images, and video and produces text, with a one-million-token context window and a maximum output of 131,072 tokens.
Outputs

What Qwen3.8-Max can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming JSON mode Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
10/10 Coding
7/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.8
Model type Multimodal
Context window 1M tokens
Maximum output 131K tokens
Release date 2026-08-03
Status Stable official release; currently available through Alibaba Cloud Model Studio
Knowledge cutoff notes

Alibaba Cloud's current model documentation and launch announcement do not provide a directly verified knowledge-cutoff date for Qwen3.8-Max. Web search and built-in tools can provide newer external information during use but do not establish the model's underlying knowledge cutoff.

Model notes

Qwen3.8-Max uses a 2.4-trillion-parameter sparse mixture-of-experts architecture with approximately 95 billion active parameters according to Alibaba's launch announcement. The model accepts text, images, and video and produces text. The official model documentation lists a maximum input length of 991,808 tokens, a one-million-token context window, a 131,072-token maximum output, and a 262,144-token maximum chain-of-thought length. Thinking mode is supported. Function calling, built-in tools including web search, structured output, context caching, and streaming are supported. The exact model page currently lists batch inference and fine-tuning as unsupported, despite broader Model Studio documentation describing batch features and batch-related pricing; the model-specific capability table is treated as authoritative here. Qwen3.8-Max-0902 is a later snapshot and should not be treated as the same canonical model record. Editorial scores are comparative estimates rather than provider-published benchmarks.

Cost

Model pricing

Input US$1.65 per 1M input tokens and US$2.00 per 1M input tokens for International scope; regional pricing varies. Cached-input pricing starts at US$0.206 per 1M tokens for the listed regional/global scope and US$0.25 per 1M tokens for International scope.
Output US$4.951 per 1M output tokens for the listed regional/global scope and US$6.00 per 1M output tokens for International scope.
Model guide

Qwen3.8-Max: Alibaba Cloud’s Million-Token Model for Long-Horizon Coding

Qwen3.8-Max is Alibaba Cloud’s flagship Qwen model for demanding multimodal reasoning, long-context analysis, coding, and agent workflows. It combines a sparse mixture-of-experts design with a one-million-token context window, image and video understanding, thinking mode, function calling, structured output, web search, and a maximum output of 131,072 tokens.

What is Qwen3.8-Max?

Qwen3.8-Max is a flagship multimodal model provided through Alibaba Cloud Model Studio. It belongs to the Qwen3.8 family and is positioned for high-complexity workloads rather than inexpensive, low-latency automation. The model is available as a stable official release through Alibaba Cloud’s Model Studio service.

Its main distinction is the combination of a very large context window and capabilities aimed at long-running tasks. Qwen3.8-Max can process text, images, and video, reason over those inputs, call functions and built-in tools, return structured output, and stream responses. Its output modality is text: despite its multimodal input support, it is not documented as a native image, video, speech, or music generation model.

Alibaba’s launch materials describe a sparse mixture-of-experts architecture with approximately 2.4 trillion total parameters and about 95 billion active parameters. In a mixture-of-experts model, different parts of the network are activated for different inputs, so the total parameter count does not mean that every parameter is used for every token. These architectural figures are provider claims, while the practical capability and score assessments below should be read separately from those claims.

Where Qwen3.8-Max fits in the Qwen lineup

Qwen3.8-Max sits at the high-capability end of Alibaba Cloud’s current Qwen offering. It is intended for complex coding, autonomous software-engineering tasks, professional document analysis, visual reasoning, long videos, and demanding research. That positioning makes it more appropriate for difficult tasks where context capacity and reasoning quality matter more than the lowest possible price or response time.

The model should not be confused with the later Qwen3.8-Max-0902 snapshot. The supplied model record treats Qwen3.8-Max as the canonical model and the later snapshot as a separate version. Applications that need reproducible behavior should verify the exact model identifier and deployment scope they use.

Context window and output limits

Qwen3.8-Max has a documented context window of up to one million tokens. The model-specific documentation lists a maximum input length of 991,808 tokens, which is slightly below the headline one-million-token context figure. In practical terms, the model can handle unusually large collections of text, long transcripts, extensive codebases, and long video-related inputs, subject to the service’s input rules and billing.

The maximum output is 131,072 tokens. The documentation also lists a maximum chain-of-thought length of 262,144 tokens when thinking mode is used. Chain-of-thought is internal reasoning generated during a thinking process; it should not automatically be treated as the same thing as the final answer returned to an application.

A large context window does not guarantee that every task will be inexpensive or equally reliable. Sending hundreds of thousands of tokens can increase input costs and may make it harder to identify the most important information. For routine prompts, a smaller and faster model may be a better operational choice.

Supported modalities and understanding

Verified model documentation describes Qwen3.8-Max as accepting:

  • Text input
  • Image input
  • Video input

Its documented output is text. This makes it suitable for asking questions about images, extracting information from video, comparing visual evidence with written instructions, or producing code and reports from multimodal material. It should not be selected when the application requires the model itself to generate an image, video, audio track, speech recording, or music file.

The model’s multimodal support is especially relevant to long-form analysis. Examples include reviewing a lengthy training recording, examining technical diagrams alongside an implementation brief, or turning visual evidence into a structured report. The research supports visual understanding and video input, but it does not establish an unrestricted duration, resolution, file-size limit, or supported codec list.

Reasoning and coding capabilities

Qwen3.8-Max supports thinking mode, which is designed for tasks that benefit from additional internal reasoning before the final response. This is useful for decomposing complex requirements, checking alternative approaches, tracing bugs, and planning multi-step work. Thinking mode can also consume more computation and may not be necessary for simple questions or straightforward extraction.

Coding is one of the model’s primary use cases. The supplied comparative editorial assessment gives it a coding score of 10 out of 10 and a reasoning score of 9 out of 10. Those scores are editorial estimates, not provider-published benchmark results. They indicate the intended evaluation of the model relative to other options in the same database, not a guaranteed performance level for every programming language or repository.

In practical terms, Qwen3.8-Max is suited to tasks such as understanding a large codebase, proposing changes across multiple files, debugging an involved failure, writing tests, reviewing implementation choices, and coordinating a sequence of software-engineering steps. Users should still inspect generated code, run tests, and control repository permissions before allowing an agent to make changes automatically.

Tools, function calling, and structured responses

Qwen3.8-Max supports function calling, built-in tools including web search, structured output, context caching, and streaming. Function calling allows an application to expose defined operations—such as querying a database or creating a ticket—so the model can request those operations in a controlled format. The model does not make an arbitrary tool call simply because it mentions one in natural language; the application must implement and authorize the available functions.

Structured output is useful when the response must conform to a defined schema rather than being a free-form paragraph. Typical examples include extracting fields from a document, returning a list of issues from a code review, or producing a machine-readable action plan. The existence of structured output does not mean that every response is automatically valid for every schema, so applications should validate returned data.

Streaming lets an application receive response content progressively instead of waiting for the complete answer. Context caching can help reduce repeated processing for reusable context, although the applicable cache rules and prices depend on Alibaba Cloud’s service configuration. Web search and other built-in tools can provide current external information during use, but tool access does not establish a fixed underlying knowledge-cutoff date. Alibaba Cloud’s supplied documentation does not provide a directly verified knowledge-cutoff date for this model.

Pricing and cost trade-offs

Alibaba Cloud’s listed prices vary by billing scope and region. The supplied pricing record gives the following figures per one million tokens:

UsageListed price
Input tokensUS$1.65 for the listed regional or global scope; US$2.00 for International scope
Cached inputFrom US$0.206 for the listed regional or global scope; US$0.25 for International scope
Output tokensUS$4.951 for the listed regional or global scope; US$6.00 for International scope

These figures should be checked against the deployment region and current Model Studio pricing before production use. The output price is materially higher than the input price, so applications can control costs by limiting unnecessary verbosity, reusing cached context where appropriate, and routing simple tasks to a less expensive model. A million-token context is valuable when the alternative is losing important information, but it is not automatically economical for ordinary prompts.

The supplied model-specific capability record marks batch inference and fine-tuning as unsupported. Broader Model Studio documentation may describe batch-related services or pricing, but the model-specific capability table is treated as authoritative for this page. Organizations that require fine-tuning or batch processing should verify whether another supported model or service is more appropriate.

Main strengths and limitations

Strengths

  • Very large context: The one-million-token context window is suited to large document collections, extensive codebases, and long-form media analysis.
  • Multimodal understanding: Text, image, and video inputs can be analyzed together, while the model returns text-based explanations, code, or structured data.
  • Complex reasoning: Thinking mode supports tasks that require planning, comparison, verification, and multiple reasoning steps.
  • Software-engineering focus: Coding, function calling, tool use, and long-horizon workflows make it suitable for advanced development assistance.
  • Application integration: Structured output, streaming, context caching, and web search support production-oriented workflows.

Limitations

  • Cost: The model is not designed for the cheapest high-volume classification or extraction workloads, particularly when prompts and outputs are large.
  • Speed: The supplied editorial speed score is 7 out of 10, suggesting a capability-versus-latency trade-off rather than a lightweight response profile. This is an editorial estimate, not a provider guarantee.
  • No native media generation: The model produces text and is not documented as an image, video, audio, speech, or music generator.
  • No fine-tuning in the model record: Fine-tuning is listed as unsupported, which may rule it out for deployments requiring custom weight adaptation.
  • Operational complexity: Tool-enabled and autonomous workflows require permission controls, validation, monitoring, and human review.
  • Unspecified knowledge cutoff: Alibaba Cloud does not provide a directly verified cutoff date in the supplied documentation. Web search can add current information during a session, but it does not change the model’s underlying training cutoff.

When to choose Qwen3.8-Max

Choose Qwen3.8-Max when the task is difficult enough to justify a high-capability model and the input is large, multimodal, or spread across several reasoning steps. It is a strong candidate for:

  • Long-horizon software-engineering agents that inspect, modify, and test a substantial codebase
  • Professional document analysis involving large collections of source material
  • Research workflows that combine long documents, visual evidence, web search, and structured findings
  • Analysis of long videos or image-heavy technical material
  • Applications that need function calling and schema-constrained responses from a reasoning-capable model
  • Tasks where preserving a large amount of context is more important than minimizing per-request cost

Another option may be more appropriate for simple classification, short summaries, routine extraction, or latency-sensitive interactive features. A smaller model can usually reduce cost and response time when the task does not require a million-token context or extended reasoning. A dedicated image, video, speech, or music generation model is the better choice when the required output is non-text media. A model with documented fine-tuning support is preferable when adapting weights to a specialized domain is a core requirement.

Bottom line

Qwen3.8-Max is best understood as a high-end text-output model for multimodal input, very large context, advanced coding, and tool-assisted reasoning. Its one-million-token context and 131,072-token maximum output make it unusually capable for long documents, long videos, and large software projects. Those advantages come with higher cost, a less lightweight speed profile, and the need for careful controls around tool use and generated code. For demanding workflows where context retention and reasoning quality matter, it is a compelling Alibaba Cloud option; for routine or media-generation tasks, a smaller or more specialized model is likely to be a better fit.


Answers to Frequently Asked Questions

How large is Qwen3.8-Max’s context window?
Qwen3.8-Max supports a headline context window of up to one million tokens, with a documented maximum input length of 991,808 tokens. Its maximum output is 131,072 tokens, while thinking mode can use up to 262,144 tokens of chain-of-thought length.
What is Qwen3.8-Max?
Qwen3.8-Max is Alibaba Cloud’s flagship multimodal model for complex workloads such as long-horizon coding, document analysis, visual reasoning, long-video analysis, and tool-assisted research. It accepts text, images, and video, but produces text output.
What can Qwen3.8-Max be used for?
The model is designed for large-scale codebase analysis, multi-file software changes, debugging, test generation, professional document analysis, long-video understanding, multimodal research, function calling, structured responses, and long-running software-engineering agents.
How much does Qwen3.8-Max cost?
The listed prices per one million tokens are US$1.65 for input, from US$0.206 for cached input, and US$4.951 for output in the listed regional or global scope. International pricing is US$2.00 for input, US$0.25 for cached input, and US$6.00 for output. Prices vary by region and should be verified against current Alibaba Cloud Model Studio pricing.
Does Qwen3.8-Max generate images, video, audio, or speech?
No. Qwen3.8-Max can understand text, images, and video, but its documented output modality is text. Applications that need image, video, audio, speech, or music generation should use a specialized generative model.


Sources 8
Provider

About Qwen