Qwen3.5

Qwen3.5-35B-A3B

by Qwen · Available; open-weight Apache 2.0 model with hosted API access through Alibaba Cloud Model Studio

A practical profile of Qwen3.5-35B-A3B covering its sparse mixture-of-experts architecture, native text-image-video input, 262K context, 65K output limit, hosted token pricing, Apache 2.0 weights, reasoning and coding uses, tool support, and deployment limitations.

Text Reasoning Coding
Qwen3.5-35B-A3B is designed to deliver strong reasoning, coding, multimodal understanding, and tool-use performance with substantially lower active compute than dense models of similar capability. It is available as downloadable Apache 2.0 weights and through Alibaba Cloud Model Studio.
Outputs

What Qwen3.5-35B-A3B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.5
Model type Multimodal
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-02-24
Status Available; open-weight Apache 2.0 model with hosted API access through Alibaba Cloud Model Studio
Knowledge cutoff notes

No authoritative first-party knowledge-cutoff date was identified for this exact model.

Model notes

The model has 35 billion total parameters with approximately 3 billion active parameters per token and uses a hybrid architecture combining Gated DeltaNet linear attention with gated full attention and sparse MoE routing. The official model card documents 256 experts with 8 routed experts plus 1 shared expert active per token. Native context is 262,144 tokens and can be extended to approximately 1,010,000 tokens with YaRN configuration, although extended-context deployment requires compatible serving configuration. Alibaba Cloud Model Studio documents text, image, and video input, text output, function calling, structured outputs, and web search support. The hosted endpoint does not support context caching, batch inference, or fine-tuning. The downloadable checkpoint is available under the Apache 2.0 license. Editorial scores are comparative estimates, not vendor-provided ratings. Qwen3.5-Flash is described by the official model card as the hosted production-oriented counterpart with additional features, but it is a separate model and is not interchangeable with this exact checkpoint.

Cost

Model pricing

Input USD 0.057 per 1 million input tokens for up to 128K input; USD 0.229 per 1 million input tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 0.25 per 1 million input tokens.
Output USD 0.459 per 1 million output tokens for up to 128K input; USD 1.835 per 1 million output tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 2 per 1 million output tokens.
Model guide

Qwen3.5-35B-A3B: A Fast, Open-Weight Multimodal MoE Model

Qwen3.5-35B-A3B is an open-weight native vision-language model from Alibaba's Qwen team with 35 billion total parameters and approximately 3 billion active parameters per token. It combines hybrid attention with sparse mixture-of-experts routing, supports text, image, and video inputs, and provides up to 262,144 tokens of native context.

What is Qwen3.5-35B-A3B?

Qwen3.5-35B-A3B is an open-weight multimodal model from Alibaba's Qwen team. It is a native vision-language model, meaning that image and video understanding are part of the model's design rather than being handled only by a separate add-on system. It accepts text, images, and video as input, then produces text responses.

The name describes its mixture-of-experts scale: the model contains approximately 35 billion total parameters, but only about 3 billion parameters are active for each token. In practical terms, this sparse design aims to provide capabilities associated with a much larger model while reducing the computation required for each generation step. The 3B figure should not be interpreted as the model's total size; local deployment still needs to accommodate the full checkpoint and its runtime requirements.

The model is provided by Alibaba's Qwen organization and is available in two main forms. Developers can download the checkpoint under the Apache 2.0 license, or use a hosted version through Alibaba Cloud Model Studio. The downloadable and hosted versions are the same model family, but deployment requirements, pricing, configuration, and operational responsibilities differ.

Architecture and context window

Qwen3.5-35B-A3B uses a hybrid architecture that combines Gated DeltaNet linear attention with gated full attention and sparse mixture-of-experts routing. Linear attention is intended to handle long sequences efficiently, while full attention is retained for interactions where detailed token-to-token relationships matter. The mixture-of-experts layer routes each token to a small subset of the available experts instead of activating every parameter.

The official model information documents 256 experts, with eight routed experts and one shared expert active per token. This architecture is a technical explanation for the model's relatively low active parameter count, not a guarantee that every workload will have the same speed or memory profile.

Native context is 262,144 tokens, or 256K tokens. That is enough for large documents, extended codebases, long conversations, and multimodal analysis where the input is within the serving system's supported limits. The model card also describes a YaRN configuration that can extend context to approximately 1,010,000 tokens. This extended range requires compatible serving configuration and should not be treated as an automatic limit of every hosted endpoint or local installation.

The maximum documented output length is 65,536 tokens. Actual usable output can still depend on the selected deployment, request settings, remaining context capacity, and service-side limits.

Inputs, outputs, and core capabilities

CapabilityQwen3.5-35B-A3B support
Text inputYes
Image inputYes
Video inputYes
Audio inputNo documented support for this model
Text outputYes
Image, video, or audio outputNo; the model is text-output only
Function and tool callingYes
Structured outputsYes through the documented hosted service
Web searchSupported through Alibaba Cloud Model Studio
StreamingSupported

Its multimodal capability is therefore primarily an understanding capability. For example, it can analyze an image, inspect video content, explain a diagram, or combine visual evidence with a long textual prompt. It does not natively generate images, video, audio, music, or speech. A workflow that needs media generation requires a separate model or service.

Reasoning and coding

Qwen3.5-35B-A3B is positioned for reasoning, coding, and tool-using agents. In a coding workflow, it can help explain existing code, propose changes, generate new code, review implementation details, and work with long technical context. Its tool-calling support also makes it suitable for applications that let a model invoke search, retrieval, databases, or other defined functions.

The model's reasoning behavior is useful for tasks such as multi-step analysis, planning, document comparison, and debugging. However, reasoning capability does not make its conclusions automatically correct. Generated code should be tested, and factual or operational decisions should be checked against authoritative sources.

Structured outputs are documented for the hosted Model Studio service. This can help applications request responses in a defined schema rather than extracting fields from free-form prose. The exact schema rules and compatibility depend on the serving interface, so developers should validate outputs in production rather than assuming that every response will satisfy a schema perfectly.

Speed, cost, and deployment trade-offs

The model's main efficiency argument is the difference between total and active parameters. With approximately 3 billion active parameters per token, it can offer a lower-compute path to multimodal reasoning than a dense model with a similar total capability target. The editorial assessment for this profile rates its speed and cost efficiency highly, but those are comparative estimates rather than scores published by Alibaba.

Local deployment offers control over data, serving configuration, and usage economics after infrastructure costs are accounted for. It also requires enough memory and compatible inference software to load and run a 35B-parameter checkpoint. The sparse routing design reduces active computation, but it does not eliminate the memory requirement associated with the complete model. The official model materials identify deployment frameworks and serving configurations, while the practical hardware requirement depends on quantization, precision, concurrency, and runtime choices.

Hosted Model Studio access avoids managing the checkpoint and infrastructure, but usage is billed by tokens and remains subject to the service's regional and operational terms. The documented endpoint does not support context caching, batch inference, or fine-tuning for this model. Those omissions matter for high-volume offline processing, repeated-prefix workloads, or teams that need to adapt the model directly to proprietary training data.

Qwen3.5-35B-A3B pricing

Alibaba Cloud's documented Global deployment pricing uses input-length tiers. For requests with up to 128K input tokens, the price is USD 0.057 per 1 million input tokens and USD 0.459 per 1 million output tokens. For requests with more than 128K and up to 256K input tokens, the price is USD 0.229 per 1 million input tokens and USD 1.835 per 1 million output tokens.

The documented international flat pricing is USD 0.25 per 1 million input tokens and USD 2 per 1 million output tokens. These are usage prices rather than a monthly subscription, and they should not be confused with the cost of running the downloadable weights locally. The applicable price depends on the selected deployment scope and current Alibaba Cloud pricing documentation.

Long-context requests can therefore cost substantially more than short-context requests, especially when they also produce long answers. Applications should avoid sending unnecessary conversation history or full documents when retrieval or targeted excerpts can provide the required context.

Main strengths and limitations

Strengths

  • Efficient sparse architecture: approximately 3B active parameters per token with 35B total parameters provides a useful capability-to-compute trade-off.
  • Native multimodal understanding: text, image, and video inputs can be handled in one model workflow.
  • Long context: the native 262,144-token window supports substantial documents, code, and multimodal prompts.
  • Open-weight availability: the Apache 2.0 checkpoint supports organizations that need more deployment control than a hosted-only model provides.
  • Agent features: tool calling, structured outputs, streaming, and hosted web search support application-oriented workflows.
  • Broad task coverage: reasoning, coding, document analysis, visual understanding, and long-context work are all central use cases.

Limitations

  • Text output only: it cannot directly generate images, video, audio, music, or speech.
  • No documented audio input: audio-centric applications need another model or a separate transcription stage.
  • Deployment size: the 35B total-parameter checkpoint may be unsuitable for small machines even though only a subset is active per token.
  • Hosted feature restrictions: the documented endpoint does not provide context caching, batch inference, or fine-tuning for this model.
  • Extended context is conditional: approximately 1.01M tokens requires YaRN and compatible serving configuration; it is not the standard native context limit.
  • Variable operational experience: local speed depends on hardware and runtime configuration, while hosted access depends on region, endpoint, pricing, and service availability.

When to choose Qwen3.5-35B-A3B

Choose Qwen3.5-35B-A3B when you need a single open-weight model for text, image, and video understanding, especially if long context, coding, reasoning, and tool use are more important than media generation. It is a strong candidate for document-analysis systems, visual question answering, code assistants, multimodal research tools, and agents that need to call external functions.

It is also a reasonable choice when the active-compute profile matters. Compared with a similarly capable dense model, its sparse mixture-of-experts design may offer a better speed or cost trade-off, although the result depends on hardware, serving software, quantization, and workload. The hosted token prices are particularly relevant for applications that keep inputs below 128K tokens; the higher long-input tier should be included in cost planning for large-document workloads.

Another option may be more appropriate when the application needs native image, video, audio, or speech generation; audio understanding; fine-tuning; batch inference; or prompt/context caching. A smaller model may also be preferable for low-memory local devices or simple high-volume tasks where multimodal reasoning is unnecessary. Conversely, users who need the extended feature set described for Alibaba's hosted production-oriented Qwen3.5-Flash should evaluate that separate model rather than assume that its capabilities are interchangeable with this checkpoint.

Bottom line

Qwen3.5-35B-A3B is a practical middle ground between compact models and much larger dense systems. Its defining combination is a 35B sparse mixture-of-experts model with roughly 3B active parameters, native text-image-video input, a 262K-token context window, open Apache 2.0 weights, and hosted tool-oriented access. Those characteristics make it particularly attractive for multimodal analysis, coding, long-context work, and local or controlled deployment.

The trade-offs are equally important: it produces text rather than media, has no documented audio input, requires substantial resources for local deployment, and lacks caching, batch inference, and fine-tuning on the documented hosted endpoint. Evaluating those constraints alongside the token pricing and deployment requirements will determine whether it is a better fit than a smaller specialized model, a media-generation model, or a more fully managed hosted alternative.


Answers to Frequently Asked Questions

What modalities and capabilities does Qwen3.5-35B-A3B support?
The model supports text, image, and video input, along with text generation, reasoning, coding, tool calling, streaming, and long-context analysis. It does not have documented audio input and cannot natively generate images, video, audio, music, or speech.
What is Qwen3.5-35B-A3B?
Qwen3.5-35B-A3B is an open-weight multimodal mixture-of-experts model from Alibaba's Qwen team. It accepts text, images, and video as input and generates text responses. The model has approximately 35 billion total parameters, with about 3 billion active per token.
How large is Qwen3.5-35B-A3B's context window?
Its native context window is 262,144 tokens, or 256K tokens. A YaRN configuration can extend the context to approximately 1.01 million tokens when the serving environment supports it. The documented maximum output length is 65,536 tokens.
How much does Qwen3.5-35B-A3B cost through Alibaba Cloud Model Studio?
For Global deployment requests with up to 128K input tokens, the documented price is USD 0.057 per 1 million input tokens and USD 0.459 per 1 million output tokens. For inputs above 128K and up to 256K tokens, the rates are USD 0.229 per 1 million input tokens and USD 1.835 per 1 million output tokens. International flat pricing is listed as USD 0.25 per 1 million input tokens and USD 2 per 1 million output tokens.
Can Qwen3.5-35B-A3B be deployed locally?
Yes. The model's checkpoint is available under the Apache 2.0 license for local deployment. However, local systems must accommodate the complete 35-billion-parameter checkpoint, even though only about 3 billion parameters are active for each token. Actual hardware requirements depend on quantization, precision, concurrency, and inference software.


Sources 6
Provider

About Qwen