Qwen3.5

Qwen3.5-397B-A17B

by Qwen · Current; open-weight model and available through Alibaba Cloud Model Studio

Qwen3.5-397B-A17B is Alibaba's Apache 2.0 open-weight native vision-language model with 397B total parameters and approximately 17B active parameters. It accepts text, images, and video, produces text, supports reasoning, coding, function calling, structured outputs, and hosted web search, and provides a 262K native context window. Its large size makes multi-GPU self-hosting necessary for practical deployment, while Model Studio provides a token-priced managed option.

Text Reasoning Coding
Qwen3.5-397B-A17B is the first large open-weight model in Alibaba's Qwen3.5 family. It is aimed at demanding multimodal tasks such as analyzing images and videos, solving complex reasoning problems, writing and reviewing code, and operating as part of tool-using agents. The model is unusually large, but its sparse mixture-of-experts design activates approximately 17 billion of its 397 billion parameters for each token. That improves efficiency compared with running every parameter on every token, although practical self-hosting still requires substantial multi-GPU infrastructure.
Outputs

What Qwen3.5-397B-A17B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.5
Model type Multimodal
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-02-16
Status Current; open-weight model and available through Alibaba Cloud Model Studio
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was found in the official model card or Alibaba Cloud model documentation.

Model notes

Canonical open-weight identifier: Qwen/Qwen3.5-397B-A17B. Hosted Alibaba Cloud Model Studio identifier: qwen3.5-397b-a17b. The model has 397B total parameters and approximately 17B active parameters per token. Native context is 262,144 tokens; the open model documentation describes extension up to approximately 1.01M tokens, but longer contexts require compatible deployment configuration. Alibaba Cloud lists structured outputs and web search as supported, while context caching, batch inference, and fine-tuning are marked unsupported for the hosted model. Open weights are released under Apache 2.0. Pricing varies by region and input-length tier.

Cost

Model pricing

Input $0.172 per 1M tokens for input up to 128K; $0.43 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $0.60 per 1M tokens.
Output $1.032 per 1M tokens for input up to 128K; $2.58 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $3.60 per 1M tokens.
Model guide

Qwen3.5-397B-A17B: Open-Weight Multimodal Reasoning at 397B Parameters

Qwen3.5-397B-A17B is Alibaba's open-weight native vision-language model with 397 billion total parameters and approximately 17 billion active parameters per token. It combines a vision encoder, hybrid gated-attention architecture, and sparse mixture-of-experts routing for text, image, and video understanding, reasoning, coding, function calling, and agent workflows. The model uses a native 262,144-token context window, produces text rather than images or other media, and is available under Apache 2.0 for self-hosting as well as through Alibaba Cloud Model Studio.

What is Qwen3.5-397B-A17B?

Qwen3.5-397B-A17B is a native vision-language model developed by Alibaba's Qwen team. In practical terms, it can process text together with images and video, then respond with text. Its intended uses include general language understanding, extended reasoning, coding, visual analysis, video understanding, graphical user interface interaction, function calling, and agent workflows.

The model has 397 billion total parameters and activates approximately 17 billion parameters for each inference step. Parameters are the learned numerical values that store a model's capabilities. Because Qwen3.5-397B-A17B uses a sparse mixture-of-experts architecture, only selected expert components process each token instead of the entire parameter set being used every time. This reduces the active computation relative to a dense 397-billion-parameter model, but it does not make the model small or easy to run locally.

Qwen3.5-397B-A17B sits at the large, capability-focused end of the current Qwen3.5 lineup. Its open-weight release is intended for developers and organizations that need more control over deployment, while Alibaba Cloud Model Studio provides a managed option for users who do not want to operate the infrastructure themselves.

Input modalities, output and core capabilities

The model accepts three input types: text, images, and video. Its native output is text. This distinction matters because the broader Qwen ecosystem includes products and models that can generate images or video, but Qwen3.5-397B-A17B itself does not natively produce image, video, or audio files.

  • Text: language understanding, generation, reasoning, summarization, and analysis.
  • Images: visual question answering and image-based reasoning.
  • Video: analysis of video content and temporally extended visual information.
  • Code: code generation, explanation, review, and analysis.
  • Tools: function calling and agent-oriented workflows.
  • Structured responses: Alibaba Cloud Model Studio lists structured outputs for the hosted version.
  • Web search: available through the Model Studio service, rather than being an inherent offline output modality of the open-weight checkpoint.

For example, a deployment could accept a screenshot and a written instruction, identify interface elements, explain what is visible, and return text or a function call for a downstream application. A video-analysis workflow could ask the model to describe events or answer questions about a recording. The supplied documentation supports these categories of use, but actual performance depends on the serving framework, prompt, input quality, and deployment configuration.

Reasoning and coding performance

Qwen3.5-397B-A17B is designed for logical and extended reasoning rather than only short conversational responses. Its large scale and multimodal design make it suitable for tasks that combine several stages of interpretation, such as examining a visual document, extracting relevant information, comparing evidence, and producing a structured answer.

Coding is another primary use case. The model can generate and analyze code, explain implementation choices, and participate in tool-driven development workflows. Function calling allows an application to expose operations that the model can request, such as retrieving data or invoking an external service. The model does not automatically make every tool call reliable, however. Applications should validate arguments, enforce permissions, and treat generated code and actions as untrusted until checked.

The research assigns editorial scores of 9 out of 10 for reasoning and coding. These are evaluation-oriented judgments for this catalog, not scores published by Alibaba and not the result of a single benchmark. They indicate that the model is positioned as a high-capability option in those areas, while real results will vary by task.

Architecture and context window

The model combines a vision encoder with a hybrid architecture built from gated delta networks and gated attention. It also uses sparse mixture-of-experts routing. The official model information specifies 512 experts, with 10 routed experts and one shared expert activated per token.

Its native context window is 262,144 tokens. A context window is the amount of input and generated conversation content the model can work with in one request, measured in tokens rather than words. This is large enough for long documents, substantial codebases, extended conversations, and multimodal analysis, subject to the limits of the particular deployment.

For the hosted Model Studio service, the documented maximum input length is 260,096 tokens and the maximum output length is 65,536 tokens. The open-weight documentation also describes a possible extension to approximately 1.01 million tokens. That larger figure should not be treated as the default: it requires a compatible serving stack and configuration, and must be validated in the intended environment. The practical baseline is therefore the native 262,144-token context.

Self-hosting and managed availability

The canonical open-weight repository is Qwen/Qwen3.5-397B-A17B on Hugging Face. The model is released under the Apache 2.0 license, which supports broad use subject to the license terms and any applicable legal or operational requirements. The documentation identifies compatibility with Transformers, vLLM, SGLang, and related serving frameworks.

Self-hosting offers control over infrastructure, deployment behavior, and data handling. It also comes with considerable costs. The repository size is approximately 807 GB, and practical full-precision or high-quality inference requires specialized multi-GPU hardware, sufficient storage, power, cooling, and operational expertise. Quantization and runtime settings may change resource requirements, but the supplied research does not establish a universal minimum hardware configuration.

Alibaba Cloud Model Studio offers a managed hosted version under the identifier qwen3.5-397b-a17b. The hosted service supports Beijing, Singapore, Frankfurt, and Virginia deployments, subject to regional availability and service-specific restrictions. Managed hosting removes much of the infrastructure burden, but introduces provider pricing, regional service dependencies, and the limits of the specific Model Studio offering.

Hosted pricing and cost trade-offs

Alibaba Cloud lists token-based pricing for Model Studio. In the Beijing, Frankfurt, and Virginia regions, requests with up to 128K input tokens are priced at $0.172 per million input tokens and $1.032 per million output tokens. For inputs above 128K and up to 256K tokens, the listed prices rise to $0.43 per million input tokens and $2.58 per million output tokens.

Singapore is listed separately at $0.60 per million input tokens and $3.60 per million output tokens. These are usage prices, not a recurring subscription fee, and the region and input-length tier materially affect the total. The input tier is based on the size of the request; output pricing applies to generated tokens.

Self-hosting avoids Alibaba's per-token inference charges, but it is not free. Hardware acquisition or rental, storage, electricity, networking, maintenance, and engineering time can outweigh hosted token costs for small or intermittent workloads. Conversely, organizations with suitable accelerator capacity and high, predictable usage may value the control and economics of operating the open-weight model themselves.

Important limitations and unsupported hosted features

The main limitation is scale. Qwen3.5-397B-A17B is not a practical choice for an ordinary laptop or a low-memory local server. Even though only part of the model is activated for each token, the full model weights and runtime memory requirements remain substantial.

The model also produces text only. It should not be selected when the requirement is native image, video, or audio generation. A separate generation model or a broader application workflow would be needed for those outputs.

For the exact hosted Model Studio model, the supplied listing marks context caching, batch inference, and fine-tuning as unsupported. These limitations apply to that hosted offering and do not necessarily prove that every third-party or self-hosted workflow is impossible. They do mean that users needing those features should verify alternatives before committing to the managed endpoint.

Long-context support should also be interpreted carefully. The native context is 262,144 tokens, while the approximately 1.01-million-token extension requires deployment-specific configuration. Longer context can increase memory and latency demands, and the research does not guarantee identical quality or availability at the extended limit.

When to choose Qwen3.5-397B-A17B

This model is a strong fit when the task justifies a large, multimodal reasoning system and the deployment team can support its infrastructure or hosted cost. Consider it for:

  • Complex image and video understanding combined with language reasoning.
  • Long-document or long-code analysis within the native context window.
  • Multimodal coding assistants and graphical interface agents.
  • Function-calling applications that need a high-capability model to select and use tools.
  • Organizations seeking an Apache 2.0 open-weight model for self-hosted experimentation or production evaluation.
  • Workloads where managed Model Studio features such as structured outputs and web search are useful.

A smaller model is likely more appropriate when low latency, low memory use, or inexpensive high-volume generation matters more than maximum capability. A model with native media generation is a better choice when the required result is an image, video, or audio file rather than a text explanation. A different hosted model or serving arrangement may also be preferable for workflows that require caching, batch inference, or fine-tuning through Model Studio.

The model's editorial speed score is 6 out of 10 and its cost score is 8 out of 10. These scores are subjective catalog assessments rather than provider-published measurements. They reflect the central trade-off: sparse activation can improve computational efficiency relative to a dense model of the same total size, but a 397-billion-parameter checkpoint still demands significantly more infrastructure than smaller alternatives.

Bottom line

Qwen3.5-397B-A17B is a large open-weight multimodal model for users who need text, image, and video understanding alongside reasoning, coding, and tool use. Its 397B total parameters, approximately 17B active parameters, 262K native context, Apache 2.0 license, and managed Model Studio availability give it both research and deployment flexibility.

Its strengths are most relevant to demanding multimodal and agentic workloads. Its drawbacks are equally clear: high infrastructure requirements, text-only output, region-dependent hosted pricing, and the absence of hosted caching, batch inference, and fine-tuning for this exact model. The best choice depends on whether the priority is maximum multimodal capability and deployment control or a smaller, faster, simpler model for routine tasks.


Answers to Frequently Asked Questions

How can Qwen3.5-397B-A17B be deployed and what hardware does it require?
The model can be self-hosted from the Qwen/Qwen3.5-397B-A17B Hugging Face repository using compatible frameworks such as Transformers, vLLM, or SGLang, or accessed through Alibaba Cloud Model Studio as qwen3.5-397b-a17b. Self-hosting requires specialized multi-GPU infrastructure, substantial storage, and significant operational resources; the repository size is approximately 807 GB.
Can Qwen3.5-397B-A17B generate images, video, or audio?
No. Qwen3.5-397B-A17B can understand text, images, and video, but its native output is text. Applications that need image, video, or audio generation must use a separate generative model or a broader workflow.
What is the context window of Qwen3.5-397B-A17B?
Qwen3.5-397B-A17B has a native context window of 262,144 tokens. The hosted Model Studio version documents up to 260,096 input tokens and 65,536 output tokens, while an approximately 1.01-million-token extension may be possible with compatible serving infrastructure and configuration.
What is Qwen3.5-397B-A17B?
Qwen3.5-397B-A17B is Alibaba’s open-weight multimodal vision-language model. It accepts text, images, and video, and produces text responses for reasoning, coding, visual analysis, video understanding, function calling, and agent workflows.
How many parameters does Qwen3.5-397B-A17B activate during inference?
The model has 397 billion total parameters and activates approximately 17 billion parameters per inference step. Its sparse mixture-of-experts architecture routes each token through selected experts, reducing active computation compared with a dense 397-billion-parameter model.


Sources 5
Provider

About Qwen