Qwen3.5

Qwen3.5-122B-A10B

by Qwen · Current and available through Alibaba Cloud Model Studio; released globally on February 24, 2026.

Alibaba Cloud’s Qwen3.5-122B-A10B is a native vision-language sparse mixture-of-experts model with 122 billion total parameters and approximately 10 billion active parameters. It accepts text, images, and video, produces text, and supports hybrid thinking, coding, function calling, structured outputs, web search, and a general 262,144-token context window. Visual endpoint limits and pricing vary by API surface and region.

Text Reasoning Coding
Qwen3.5-122B-A10B is a large multimodal model released globally on February 24, 2026. It combines text reasoning, coding, image understanding, and video understanding with a sparse mixture-of-experts design intended to reduce the amount of computation used for each request. The model sits below Qwen3.5-397B-A17B in Alibaba’s Qwen3.5 lineup and is aimed at applications that need advanced visual analysis without using the largest model in the family.
Outputs

What Qwen3.5-122B-A10B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Structured output
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.5
Model type Multimodal
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-02-24
Status Current and available through Alibaba Cloud Model Studio; released globally on February 24, 2026.
Knowledge cutoff notes

Alibaba's current model documentation and model card reviewed for this record do not provide a directly verified knowledge-cutoff date for the exact Qwen3.5-122B-A10B model.

Model notes

Qwen3.5-122B-A10B is a native vision-language sparse mixture-of-experts model with 122B total parameters and approximately 10B active parameters. Alibaba describes its architecture as combining linear attention with sparse MoE routing. The general Model Studio API documents a 262,144-token context window, 260,096 maximum input tokens, and 65,536 maximum output tokens. A separate visual-understanding table lists a 32K context and 8K maximum output for the visual endpoint, so limits can depend on the API surface and request mode. The model supports hybrid thinking, function calling, structured outputs, and web search in supported Model Studio deployment scopes. Context caching, batch inference, and fine-tuning are documented as unsupported. Editorial scores are comparative estimates rather than vendor benchmarks.

Cost

Model pricing

Input $0.115 per 1M tokens for input up to 128K; $0.287 per 1M tokens for input from 128K to 256K. Pricing varies by deployment scope and region.
Output $0.917 per 1M tokens for requests up to 128K input; $2.294 per 1M tokens for requests from 128K to 256K input. Pricing varies by deployment scope and region.
Model guide

Qwen3.5-122B-A10B: Alibaba’s Sparse MoE Model for Visual Reasoning

Qwen3.5-122B-A10B is Alibaba’s native vision-language mixture-of-experts model with 122 billion total parameters and approximately 10 billion active parameters. It accepts text, images, and video, produces text, and supports hybrid thinking, function calling, structured outputs, and web search through supported Alibaba Cloud Model Studio deployments.

What is Qwen3.5-122B-A10B?

Qwen3.5-122B-A10B is a multimodal, native vision-language model from Alibaba Cloud. It can process text, images, and video in the same general model family, while its documented output is text. That makes it suitable for tasks such as explaining an image, extracting information from a document, interpreting a chart, answering questions about video content, writing code, or combining visual evidence with a longer reasoning task.

The name describes its scale and sparse architecture. The model has 122 billion total parameters, but approximately 10 billion active parameters are used for a particular input through mixture-of-experts routing. In simple terms, the model contains many specialist processing components, while a routing mechanism selects a smaller subset for each request. Alibaba describes the architecture as combining linear attention with sparse MoE routing. This is intended to provide a balance between large-model capability and more manageable inference costs, although actual speed and cost depend on the deployment scope, region, request size, and API surface.

Qwen3.5-122B-A10B was released globally on February 24, 2026, and is currently listed through Alibaba Cloud Model Studio. It is positioned below Qwen3.5-397B-A17B, the larger sibling referenced in Alibaba’s Qwen3.5 lineup. The model is therefore best understood as a high-capability general-purpose multimodal option rather than a small, low-latency model.

Capabilities and supported modalities

The model accepts three documented input types: text, images, and video. It produces text rather than images, audio, or video. It should not be selected for native media generation. Its visual capabilities are focused on understanding and reasoning about supplied media.

  • Text input and output: Suitable for conversation, analysis, summarization, generation, and coding.
  • Image input: Useful for image interpretation, document analysis, chart reading, and visual question answering.
  • Video input: Supports analysis of video content, subject to the limits of the specific visual endpoint or deployment.
  • Audio: Audio input and audio output are not documented for this model.
  • Image, video, and audio generation: Not supported as native model outputs.

Alibaba documents hybrid thinking support. This allows the model to work in a mode that can expose or use additional reasoning before producing an answer, depending on the API configuration. Reasoning support does not guarantee correctness: visual interpretation, long documents, ambiguous instructions, and factual claims still require appropriate validation.

Context and output limits

The general Model Studio API documentation lists a context window of 262,144 tokens, with a maximum of 260,096 input tokens and 65,536 output tokens. A token is a unit of text or other encoded content; the context window is the total amount of information the model can handle in one request and response. These limits make the model suitable for long documents, extended coding tasks, and workflows that combine substantial text with visual material.

There is an important endpoint-specific qualification. Alibaba’s separate visual-understanding table lists a 32K context limit and an 8K maximum output for the visual endpoint. The larger 262,144-token and 65,536-token limits therefore should not automatically be assumed for every image or video request. Before production deployment, verify the limits for the exact Model Studio endpoint, region, and request mode being used.

The large general context limit is useful for applications such as reviewing a lengthy technical document alongside charts, examining a substantial codebase, or maintaining a long agent workflow. It does not mean that every request will be equally fast or equally inexpensive. Larger inputs consume more tokens and can increase latency and cost.

Reasoning, coding, and tool support

Qwen3.5-122B-A10B is designed for advanced reasoning across text and visual inputs. Practical examples include asking it to compare figures in a chart, identify inconsistencies between a written report and an image, explain a diagram, or reason about a sequence of events in a video. These are model capabilities supported by the documented multimodal design; they should not be confused with a guarantee of reliable perception in every case.

The model also supports coding use cases. It can generate, explain, transform, and review code through text interaction, and its long context can help with larger code or documentation inputs. The supplied evaluation assigns a comparative editorial coding score of 9 out of 10 and a reasoning score of 9 out of 10. Those scores are editorial estimates, not Alibaba-published benchmark results, so they should be treated as directional rather than as verified performance measurements.

Function calling is supported, allowing an application to provide tools that the model can request during a conversation. For example, an application could expose a database lookup, an internal search function, or a document retrieval operation. The model can also produce structured outputs, which is useful when an application needs predictable fields rather than free-form prose. These features do not mean that the model independently performs arbitrary external actions; the surrounding application must define, validate, and execute the available tools.

Supported Model Studio deployments also document web search. Web search can provide a route to web-grounded responses, but the result depends on the configured search capability and should still be checked when accuracy or freshness matters.

Pricing and deployment considerations

Alibaba Cloud Model Studio pricing is token-based and varies by deployment scope and region. The supplied pricing information lists the following rates:

Request categoryInput priceOutput price
Requests with up to 128K input tokens$0.115 per 1 million tokens$0.917 per 1 million tokens
Requests with 128K to 256K input tokens$0.287 per 1 million tokens$2.294 per 1 million tokens

The output rate is determined by the input-length category in the supplied pricing table, so a request using more than 128K input tokens can cost more for both input and output tokens. These figures are not a flat subscription price, and they do not necessarily apply identically to every deployment scope or region. Confirm the current Model Studio price page before estimating production spend.

The model’s sparse MoE architecture may offer a better capability-to-computation trade-off than a dense model with a similar total parameter count, but the available research does not establish a universal latency or cost advantage. The editorial speed score is 8 out of 10 and the cost score is 8 out of 10; both are comparative estimates rather than provider measurements. Requests involving video, long contexts, tool calls, or extended reasoning may have different practical performance from short text prompts.

Main strengths and limitations

Strengths

  • Multimodal reasoning: It combines text, image, and video understanding in one model rather than requiring a separate vision model for every workflow.
  • Large documented context: The general API supports up to 262,144 tokens, subject to endpoint-specific restrictions.
  • Sparse architecture: Approximately 10 billion parameters are active from a 122-billion-parameter total, providing a large-model design with selective routing.
  • Developer features: Function calling, structured outputs, and supported web search enable agentic and application-oriented workflows.
  • Coding and analysis: The model is suited to code generation, document analysis, chart interpretation, and visual question answering.

Limitations

  • Text-only output: It does not natively generate images, audio, or video.
  • Endpoint-dependent limits: The visual-understanding endpoint is documented with a 32K context and 8K maximum output, which is substantially lower than the general API limits.
  • Not a small model: Applications requiring the lowest latency, a minimal hosting footprint, or simple low-cost text completion may be better served by a smaller option.
  • No documented fine-tuning or batch support: The supplied Model Studio research lists fine-tuning, batch inference, and context caching as unsupported.
  • Pricing varies: Region and deployment scope affect pricing, and long requests move into a higher price tier.
  • Visual reasoning still needs review: The model can analyze images and video, but its outputs should be checked when decisions depend on precise visual interpretation.

When to choose Qwen3.5-122B-A10B

Choose Qwen3.5-122B-A10B when one workflow needs substantial reasoning across text and visual material. It is a strong candidate for document and chart analysis, video understanding, multimodal research assistants, coding agents that inspect screenshots or diagrams, and tool-using systems that need structured responses. Its combination of a large general context, visual inputs, function calling, and web search is particularly useful when a task involves several information sources rather than a single short prompt.

It is also a sensible middle position in the Qwen3.5 lineup when the larger Qwen3.5-397B-A17B would be excessive for the workload. The supplied research supports that lineup positioning, but it does not provide a direct benchmark comparison between the two models. The choice should therefore be based on the required quality, budget, latency, and deployment availability rather than assuming a fixed performance difference.

Consider another type of model when the task is narrowly defined. A smaller text-only model may be more appropriate for high-volume classification, short autocomplete requests, or latency-sensitive applications. A dedicated image, video, audio, or media-generation model may be preferable when the required output is not text. A model or service with documented fine-tuning, batch inference, or caching support may better fit those operational requirements.

Bottom line

Qwen3.5-122B-A10B is a large sparse MoE model for applications that need advanced text reasoning together with image and video understanding. Its approximately 10 billion active parameters, large general API context, structured outputs, function calling, and supported web search make it more suitable for complex multimodal workflows than for simple text generation. The main checks before adoption are the exact endpoint limits, regional pricing, expected latency, and whether text-only output is sufficient for the application.


Answers to Frequently Asked Questions

When should you choose Qwen3.5-122B-A10B?
Choose it for complex multimodal workflows involving substantial reasoning across text, images, and video, such as document and chart analysis, video understanding, coding agents that inspect diagrams or screenshots, and tool-using assistants. A smaller text-only model may be better for simple, high-volume, or latency-sensitive tasks.
How much does Qwen3.5-122B-A10B cost?
The supplied Model Studio pricing lists $0.115 per 1 million input tokens and $0.917 per 1 million output tokens for requests with up to 128K input tokens. For requests with 128K to 256K input tokens, the rates are $0.287 per 1 million input tokens and $2.294 per 1 million output tokens. Actual pricing may vary by deployment scope and region.
What are the context and output limits of Qwen3.5-122B-A10B?
The general Model Studio API documents a 262,144-token context window, with up to 260,096 input tokens and 65,536 output tokens. However, Alibaba’s visual-understanding endpoint is listed with a 32K context limit and an 8K maximum output, so the exact limits depend on the endpoint, region, and request mode.
What is Qwen3.5-122B-A10B?
Qwen3.5-122B-A10B is a multimodal, native vision-language model from Alibaba Cloud that processes text, images, and video while producing text output. It uses a sparse mixture-of-experts architecture with 122 billion total parameters and approximately 10 billion active parameters per request.
What modalities does Qwen3.5-122B-A10B support?
The model accepts text, images, and video as inputs and returns text. It can analyze documents, charts, images, and video content, but it does not natively generate images, video, audio, or other media, and audio input and output are not documented.


Sources 6
Provider

About Qwen