Qwen3.6

Qwen3.6-Max-Preview

by Qwen · Preview; currently accessible; scheduled for deprecation on October 10, 2026

Qwen3.6-Max-Preview is Alibaba Cloud’s preview Max model for advanced coding agents, front-end development, long-context analysis, function calling, structured outputs, and web-search-enabled text generation. It offers a 262,144-token context window and up to 65,536 output tokens, with international input pricing from $1.30 to $2.00 per million tokens and output pricing from $7.80 to $12.00 per million tokens depending on request length. The text-only model does not support batch inference or fine-tuning and is scheduled for deprecation on October 10, 2026.

Text Reasoning Coding
Qwen3.6-Max-Preview is Alibaba Cloud’s preview release of the largest model in the Qwen3.6 family. It targets demanding text-generation tasks such as coding-agent workflows, front-end development, long-context analysis, function calling, and structured API responses. The model supports a 262,144-token context window and up to 65,536 output tokens, but its preview status and scheduled retirement make lifecycle planning especially important for anyone considering a new integration.
Outputs

What Qwen3.6-Max-Preview can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
7/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.6
Model type General Purpose
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-04-20
Status Preview; currently accessible; scheduled for deprecation on October 10, 2026
Deprecation date 2026-10-10
Shutdown date 2026-10-10
Knowledge cutoff notes

Alibaba Cloud's current model documentation does not provide a directly verifiable knowledge-cutoff date for this exact model. Web search support is a separate provider integration and does not establish the underlying knowledge cutoff.

Model notes

The canonical Alibaba Cloud Model Studio model ID is qwen3.6-max-preview. It is a text-only preview model with function calling, structured outputs, web search, prefix completion, and context caching. The standard context window is 262,144 tokens, with 245,760 maximum input tokens and 65,536 maximum output tokens. Thinking mode supports 229,376 input tokens, 65,536 output tokens, and a 131,072-token maximum chain-of-thought length. Batch inference and fine-tuning are listed as unsupported. As of October 7, 2026, the model remains accessible but is scheduled for deprecation on October 10, 2026; Alibaba Cloud identifies qwen3.7-max as the replacement. Editorial scores are comparative estimates, not vendor-published ratings.

Cost

Model pricing

Input $1.30 per 1M tokens for input up to 128K; $2.00 per 1M tokens for input above 128K and up to 256K, international pricing
Output $7.80 per 1M tokens for requests up to 128K input; $12.00 per 1M tokens for requests above 128K and up to 256K input, international pricing
Model guide

Qwen3.6-Max-Preview: Alibaba Cloud’s Long-Context Coding Model Nearing Deprecation

Qwen3.6-Max-Preview is Alibaba Cloud Model Studio’s preview-stage Max model in the Qwen3.6 family. It is designed for advanced text generation, coding agents, front-end development, function calling, structured outputs, web search, and long-context workloads. Its 262,144-token context window and 65,536-token maximum output make it suitable for substantial prompts and extended responses, but the model is text-only, does not support batch inference or fine-tuning, and is scheduled for deprecation on October 10, 2026.

What is Qwen3.6-Max-Preview?

Qwen3.6-Max-Preview is a preview-stage text-generation model available through Alibaba Cloud Model Studio. It is positioned as the largest and most capable member of the Qwen3.6 family, with an emphasis on advanced coding, coding-agent execution, front-end development, long-context work, and retaining information across lengthy prompts.

The model accepts text and produces text. Its name is important for two reasons: “Max” identifies its position at the high-capability end of the Qwen3.6 lineup, while “Preview” indicates that it is not presented as a stable, long-term model target. Alibaba Cloud lists it as accessible during the preview period but schedules it for deprecation on October 10, 2026. Alibaba Cloud identifies Qwen3.7-Max as the replacement.

Where it fits in Alibaba Cloud’s lineup

Qwen3.6-Max-Preview is intended for workloads that need more capability and context than smaller or less advanced model options in the Qwen family. Alibaba Cloud specifically positions it for stronger coding-agent execution, “vibe coding,” front-end development, and long-tail knowledge retention than earlier Qwen3-Max and Qwen3.6-Plus models.

That positioning does not make it the best choice for every request. A high-capability model can be useful when a task involves large repositories, complicated instructions, multi-step reasoning, or substantial generated code. For short, routine, or cost-sensitive requests, a smaller model may be more economical. The supplied research does not provide a direct price or benchmark comparison with those sibling models, so these are practical trade-offs rather than quantified performance claims.

Main capabilities

Qwen3.6-Max-Preview combines ordinary text generation with features that make it useful inside applications:

  • Advanced coding: It is intended for coding assistance, coding agents, and front-end development, including workflows sometimes described as vibe coding.
  • Function calling: Applications can allow the model to request registered functions or tools. The application, rather than the model alone, performs the external action and returns the result.
  • Structured outputs: The model can produce data in a structured format for application workflows, reducing the need to extract fields from free-form prose.
  • Web search: Alibaba Cloud supports web search through its provider integration. This adds a way to retrieve current information, but it does not change the model’s underlying knowledge cutoff.
  • Prefix completion: The model supports continuing text from an existing prefix, which can be useful in certain generation and coding scenarios.
  • Context caching: Reusable prompt content can be cached to reduce the cost or processing burden of repeated context, subject to Alibaba Cloud’s caching rules and prices.
  • Streaming: Responses can be delivered incrementally rather than waiting for the complete response.

These features make the model suitable for applications that combine language generation with tools, APIs, or structured data handling. They do not mean that Qwen3.6-Max-Preview can independently browse the web or execute arbitrary actions without the surrounding Model Studio integration and application logic.

Context window and output limits

The model has a maximum context window of 262,144 tokens. A token is a unit of text used by the model; the count includes the input and the generated response within the request’s overall limits. This large window is useful for source-code repositories, long specifications, transcripts, document collections, and multi-stage agent sessions.

Alibaba Cloud lists a standard maximum input length of 245,760 tokens and a maximum output length of 65,536 tokens. In thinking mode, the maximum input length is 229,376 tokens, while the maximum output remains 65,536 tokens. Thinking mode also has a maximum chain-of-thought length of 131,072 tokens.

These limits should not be interpreted as a requirement to send very large prompts. Large requests can increase latency and cost, and relevant information still needs to be organized clearly. For many applications, retrieving only the most relevant files or documents will be more efficient than sending an entire data collection on every request.

Pricing and context caching

For international Model Studio pricing, requests with up to 128,000 input tokens cost $1.30 per million input tokens and $7.80 per million output tokens. For requests with more than 128,000 and up to 256,000 input tokens, the prices increase to $2.00 per million input tokens and $12.00 per million output tokens.

Request input lengthInput priceOutput price
Up to 128,000 tokens$1.30 per million tokens$7.80 per million tokens
Above 128,000 up to 256,000 tokens$2.00 per million tokens$12.00 per million tokens

Alibaba Cloud also publishes regional prices in Chinese yuan for China and Singapore deployments. The provider’s explicit cache-creation and cache-read prices are lower than standard input pricing, although the exact amount depends on the applicable pricing region and cache operation. Teams repeatedly sending the same large instructions or reference material should examine caching economics before assuming that the full context cost will be paid on every request.

The output price is notably important for planning. Long responses, extensive code generation, and extended reasoning can cost substantially more than short answers, particularly when the request falls into the higher input-length tier. The model’s editorial cost score is 5 out of 10; this is a comparative editorial estimate, not a rating published by Alibaba Cloud.

Reasoning and coding suitability

Qwen3.6-Max-Preview supports a thinking mode with a maximum chain-of-thought length of 131,072 tokens. This gives it room for complex multi-step processing, although the supplied research does not include a standardized reasoning benchmark or a provider-published reasoning score.

Its strongest practical fit is coding that requires more than simple autocomplete. Examples include planning changes across multiple files, analyzing a large codebase, generating front-end components, coordinating tool calls, and iterating through an agent workflow. The editorial coding score is 9 out of 10, while the editorial reasoning score is 8 out of 10. These scores are subjective comparative assessments and should not be treated as benchmark results.

The model’s editorial speed score is 7 out of 10. Large contexts, thinking mode, long outputs, and tool interactions can all increase response time. If an application needs very fast responses to small requests, a smaller or more speed-focused option may be preferable.

Modalities and unsupported features

Qwen3.6-Max-Preview is text-only. It accepts text input and returns text output; it does not directly generate images, audio, or video, and the supplied model data does not list image, audio, or video input support. Applications requiring multimodal understanding or non-text generation should use an option designed for those modalities instead.

The current Model Studio listing also marks batch inference and fine-tuning as unsupported. This limits its suitability for large offline processing pipelines and for organizations that need to adapt the model with their own training data. Function calling, structured outputs, web search, streaming, and context caching are supported, but these application features should not be confused with fine-tuning or independent model customization.

Best use cases

  • Advanced coding assistants that need to inspect substantial code or specifications.
  • Coding-agent workflows involving planning, tool calls, and multiple implementation steps.
  • Front-end development and generation of larger interface components.
  • Long-context analysis of technical documents, repositories, or extended conversations.
  • Structured API workflows where the application needs predictable fields in the response.
  • Text-generation tasks that benefit from provider-supported web search.
  • Applications that repeatedly reuse large instructions or reference material and can benefit from context caching.

When to choose this model

Choose Qwen3.6-Max-Preview when the task justifies a high-capability text model with a very large context window. It is a reasonable candidate for a coding agent that must work across many files, a front-end development assistant, or a document-analysis system that needs to keep a large amount of context available. Function calling and structured outputs also make it more suitable than a plain text-completion model for applications that connect language generation to business logic.

Its preview status changes the decision. A new production system that must remain stable for a long period should account for the October 10, 2026 deprecation date and migration to Qwen3.7-Max. If the integration is exploratory, temporary, or easy to migrate, the model may still be useful while it remains available.

Consider another option when the main priority is low cost, rapid responses, batch processing, fine-tuning, or image, audio, or video support. Qwen3.6-Max-Preview’s high output pricing and unsupported batch and fine-tuning features can outweigh its capability advantages for those workloads. Similarly, a smaller model may be a better fit for short classification, extraction, or routine chat requests that do not require a 262,144-token context window.

Lifecycle and final assessment

Alibaba Cloud lists Qwen3.6-Max-Preview as released on April 20, 2026, and scheduled for deprecation on October 10, 2026. As of October 7, 2026, the supplied research indicates that it remains accessible. The short remaining lifecycle is therefore a central part of the model’s evaluation, not a minor administrative detail.

Technically, the model is aimed at demanding text workloads: it offers a 262,144-token context window, up to 65,536 output tokens, thinking mode, coding-agent support, function calling, structured outputs, web search, streaming, and caching. Practically, those benefits must be balanced against higher costs for long requests and outputs, text-only operation, the absence of batch inference and fine-tuning, and imminent deprecation. It is best viewed as a capable preview model for advanced experimentation or workloads with a clear migration plan rather than an uncomplicated long-term default.


Answers to Frequently Asked Questions

When will Qwen3.6-Max-Preview be deprecated, and what will replace it?
Alibaba Cloud has scheduled Qwen3.6-Max-Preview for deprecation on October 10, 2026. Qwen3.7-Max is identified as its replacement, so production systems using the preview model should plan for migration.
What are the best use cases for Qwen3.6-Max-Preview?
The model is well suited to advanced coding assistants, coding agents that plan and execute multi-step tasks, front-end development, large repository analysis, long technical documents, structured API workflows, and applications that reuse large prompts through context caching.
How much does Qwen3.6-Max-Preview cost?
For international Model Studio pricing, requests with up to 128,000 input tokens cost $1.30 per million input tokens and $7.80 per million output tokens. Requests above 128,000 and up to 256,000 input tokens cost $2.00 per million input tokens and $12.00 per million output tokens. Alibaba Cloud also offers regional pricing and separate cache-creation and cache-read pricing.
What is the context window of Qwen3.6-Max-Preview?
Qwen3.6-Max-Preview supports a maximum context window of 262,144 tokens. Alibaba Cloud lists a maximum input length of 245,760 tokens and a maximum output length of 65,536 tokens in standard mode. In thinking mode, the maximum input length is 229,376 tokens, with a maximum chain-of-thought length of 131,072 tokens.
What is Qwen3.6-Max-Preview?
Qwen3.6-Max-Preview is a preview-stage text-generation model available through Alibaba Cloud Model Studio. It is designed for advanced coding, coding-agent workflows, front-end development, long-context analysis, function calling, structured outputs, web search, streaming, and context caching.


Sources 6
Provider

About Qwen