Qwen3-VL

Qwen3-VL-235B-A22B-Instruct

by Qwen · Current and accessible; open-weight release and Alibaba Cloud Model Studio API availability

Alibaba Cloud's Qwen3-VL-235B-A22B-Instruct is a large mixture-of-experts vision-language model with approximately 235B total and 22B active parameters. It accepts text, images and video, returns text, and targets advanced OCR, document analysis, video comprehension, visual coding, spatial reasoning and visual-agent workflows. The hosted endpoint provides a 131,072-token context window and regional API pricing, while the open-weight release supports self-hosted deployment under Apache 2.0.

Text Reasoning Coding
Qwen3-VL-235B-A22B-Instruct is the instruction-following model in Alibaba Cloud's Qwen3-VL family. Released on September 23, 2025, it is designed for applications that need detailed visual interpretation alongside ordinary language generation. It can analyze images and video, read difficult documents, reason about spatial relationships and produce text-based answers or code. Its open-weight release provides a path to self-hosting, while Alibaba Cloud Model Studio offers a managed API with region-dependent features and pricing.
Outputs

What Qwen3-VL-235B-A22B-Instruct can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Streaming Structured output
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3-VL
Model type Multimodal
Context window 131K tokens
Maximum output 33K tokens
Release date 2025-09-23
Status Current and accessible; open-weight release and Alibaba Cloud Model Studio API availability
Knowledge cutoff notes

No authoritative first-party knowledge-cutoff date was identified for this exact model.

Model notes

The exact hosted model ID is qwen3-vl-235b-a22b-instruct. Alibaba Cloud lists text, image and video inputs with text output. Function calling is supported in the China Beijing deployment but not listed as supported in Singapore, Germany or US Virginia. Structured outputs are supported, but the current model documentation does not separately confirm legacy JSON mode. The hosted endpoint lists a 131,072-token context window, while the open-weight documentation describes a native 256K configuration and optional YaRN extension to 1M tokens in supported self-hosted deployments. The open-weight repository reports approximately 236B parameters and BF16 weights; the model name identifies 235B total parameters and 22B active parameters. Pricing varies by region and promotions.

Cost

Model pricing

Input $0.287 per 1 million tokens in the US Virginia global deployment; $0.400 per 1 million tokens in Singapore
Output $1.147 per 1 million tokens in the US Virginia global deployment; $1.600 per 1 million tokens in Singapore
Model guide

Qwen3-VL-235B-A22B-Instruct: Open-Weight Vision-Language Model for OCR, Video and Visual Reasoning

Qwen3-VL-235B-A22B-Instruct is Alibaba Cloud's large open-weight mixture-of-experts vision-language model for text, image and video understanding. It combines 235 billion total parameters with approximately 22 billion active parameters and is aimed at demanding workloads such as document analysis, advanced OCR, long-video comprehension, visual coding, spatial reasoning and visual-agent workflows. The hosted Model Studio endpoint accepts text, images and video, returns text, provides a 131,072-token context window and supports structured outputs.

What Qwen3-VL-235B-A22B-Instruct is

Qwen3-VL-235B-A22B-Instruct is a large vision-language model from Alibaba Cloud's Qwen team. A vision-language model processes visual information together with text, allowing an application to ask questions about an image, extract information from a document, interpret a video or connect visual evidence with a written response.

The model is an instruction-following variant, which means it is intended for direct task completion and conversational applications. It is separate from the Qwen3-VL Thinking variant, which is positioned for a different reasoning style. The current model's primary output is text: it can describe, analyze, classify or transform multimodal input, but it does not natively return generated images, video, audio or music.

Qwen3-VL-235B-A22B-Instruct uses a mixture-of-experts architecture. It has approximately 235 billion total parameters, with about 22 billion active parameters used for an individual inference according to the model's naming and supplied documentation. This design gives it a large overall capacity while avoiding activation of every parameter for every request. It remains a demanding model to operate, particularly in self-hosted deployments.

What the model can understand and do

The model is built for high-end multimodal understanding rather than a narrow image-captioning task. Alibaba's documentation and the open-weight model materials emphasize several related capabilities.

  • Image understanding: It can interpret image content, answer visual questions, describe scenes and perform fine-grained visual analysis.
  • OCR and document intelligence: It is designed to read difficult documents, including blurred, tilted, low-light and multilingual material. Potential applications include forms, invoices, charts and other image-based records.
  • Video understanding: It can analyze video over time, identify events and support timestamp-aware interpretation. This makes it more suitable for reviewing recorded material than a model limited to a single still image.
  • Visual coding: It can use visual references to help produce HTML, CSS, JavaScript and diagram-oriented formats. For example, a developer could provide a screenshot or design reference and ask for a text-based implementation.
  • Spatial reasoning: The model is designed to reason about object positions, viewpoints, occlusion and two-dimensional grounding, with selected three-dimensional grounding scenarios described in the model materials.
  • Visual-agent workflows: When connected to an appropriate execution system, the model family can recognize interface elements, infer their functions, call tools and assist with computer or mobile interface tasks.

These are model and provider capability descriptions, not a guarantee that every application will perform equally well. Results depend on the input quality, prompt, serving configuration, region and the surrounding software used to execute tools or process files.

Context window and output limits

The Alibaba Cloud Model Studio endpoint lists a 131,072-token context window. Its documented maximum input length is 129,024 tokens and its maximum output length is 32,768 tokens. A token is a unit of text used by the model; it is not identical to a word, so the practical amount of readable content varies by language and document format.

These hosted limits should be distinguished from self-hosted configuration claims. The open-weight documentation describes a native 256K-token configuration and provides YaRN guidance for extending supported deployments to as much as 1 million tokens. Those figures describe possible open-weight serving configurations, not the default limit of the managed Model Studio endpoint. Hardware, inference software and configuration determine whether an extended context is practical.

The large context is useful for long documents, collections of pages, lengthy transcripts and video-related prompts. However, a large theoretical limit does not remove the need to manage input quality and relevance. Sending unnecessary material increases processing cost and can make it harder for an application to obtain a focused answer.

Hosted API and open-weight deployment

Alibaba Cloud Model Studio provides the hosted model ID qwen3-vl-235b-a22b-instruct. The documented endpoint accepts text, images and video, returns text and supports structured outputs. Structured output can help an application request responses that follow a specified data shape, although the supplied documentation does not separately confirm a legacy JSON-mode capability.

Function calling is deployment-dependent. The supplied model documentation lists it as supported in the China Beijing deployment, but does not list it as supported for Singapore, Germany or US Virginia. Developers should therefore verify the region-specific feature matrix before designing a production workflow around tool calls.

The hosted endpoint does not list web search, context caching, batch inference or fine-tuning as supported capabilities for this exact model. The model's visual-agent positioning should not be interpreted as built-in autonomous computer control: tool execution and interface actions require an external system that receives and carries out the model's instructions.

The open-weight release is published under the Apache 2.0 license on Hugging Face. Official deployment guidance covers Transformers, vLLM and SGLang. Self-hosting offers more control over configuration and data handling, but the scale of the model makes it a substantial multi-GPU workload. Quantized or FP8 variants can reduce memory requirements, but they do not make the model a lightweight option for ordinary consumer hardware.

API pricing

Alibaba Cloud lists the original price for the US Virginia global deployment at $0.287 per 1 million input tokens and $1.147 per 1 million output tokens. The supplied research reports the same original rates for Germany and China Beijing. Singapore is listed at $0.400 per 1 million input tokens and $1.600 per 1 million output tokens.

These are token-based API prices rather than a consumer subscription fee. Actual charges can vary by region, promotional offer and the amount of input and output generated. Self-hosting replaces per-token API charges with infrastructure, storage, engineering and operational costs. For occasional workloads, the hosted endpoint may be simpler; for sustained high-volume use, the economics require a comparison between API spending and the cost of operating suitable hardware.

Strengths and limitations

Strengths

  • Broad visual input: Text, image and video support are available through the documented hosted endpoint.
  • Strong fit for document-heavy work: OCR, multilingual document reading and difficult visual layouts are central use cases.
  • Video-oriented reasoning: Temporal interpretation and timestamp-aware analysis extend its usefulness beyond still-image question answering.
  • Visual coding and spatial analysis: The model is aimed at screenshots, diagrams, interfaces and other visually grounded technical tasks.
  • Open-weight availability: The Apache 2.0 release allows organizations to evaluate self-hosted deployment options rather than relying only on a managed endpoint.
  • Large context options: The managed endpoint provides 131,072 tokens, while supported self-hosted configurations offer larger documented possibilities.
  • Structured responses: Structured outputs are listed for the hosted model, which can make integration with downstream software more predictable.

Limitations

  • High operating cost and complexity: A model of this size is difficult to self-host and is not a natural choice for low-resource deployments.
  • Different hosted and self-hosted limits: The 131,072-token Model Studio limit should not be confused with the 256K native configuration or optional 1M extension described for open-weight serving.
  • Regional feature differences: Function calling availability depends on the deployment region.
  • Limited hosted workflow features: The exact endpoint does not list web search, caching, batch inference or fine-tuning.
  • Text-only output: It understands images and video but does not directly generate image, video or audio files.
  • Latency trade-off: Its large architecture and multimodal processing make it better suited to quality-sensitive analysis than to applications that require consistently minimal response times.

Editorially, the model's strongest value is its combination of high-end visual understanding, open-weight availability and broad multimodal coverage. That assessment is not a provider-published benchmark score; it reflects the documented capabilities and the practical trade-off between capability, speed, infrastructure and price.

Reasoning, coding and tool support

Qwen3-VL-235B-A22B-Instruct is intended to reason over visual evidence, not merely identify objects. Examples include locating an item in a scene, interpreting a diagram, following events through a video or connecting a document field with a question. The supplied editorial assessment rates its reasoning capability highly, but no benchmark result is provided here, so that assessment should not be treated as an official score.

Its coding usefulness is especially tied to visual inputs. It can help translate screenshots and visual references into HTML, CSS, JavaScript or diagrams, and it can support interface understanding. It produces code as text; the supplied research does not confirm built-in code execution. Any generated code should therefore be reviewed and tested in a separate environment.

Tool use is available only where the serving deployment exposes the relevant function-calling capability. In China Beijing, the documentation lists function calling for this model. The supplied regional documentation does not list that capability for the US Virginia, Germany or Singapore deployments. Applications that need tools in another region should verify current documentation rather than assuming that the model family feature is universally enabled.

Best use cases

This model is a strong candidate for applications where visual detail and broad context matter more than minimum latency or simple infrastructure. Suitable workloads include:

  • Extracting fields from invoices, forms, reports and other difficult documents.
  • Analyzing charts, diagrams, screenshots and technical illustrations.
  • Reviewing long videos for events, scenes or time-specific information.
  • Building image-grounded question-answering and visual search assistants.
  • Converting interface screenshots or visual designs into initial web code.
  • Supporting UI understanding and multimodal agent systems with an external execution layer.
  • Processing long multimodal records where a smaller context window would require aggressive document splitting.

When to choose this model

Choose Qwen3-VL-235B-A22B-Instruct when the workload needs one model to handle text, images and video, and when OCR, spatial interpretation, visual coding or long-context analysis are more important than the lowest possible cost and latency. The open-weight release is particularly relevant to teams that need deployment control and have the infrastructure to operate a very large model. The managed API is more practical for teams that want access without building a multi-GPU serving stack.

A smaller Qwen3-VL model is likely more appropriate when response speed, hardware cost or operational simplicity is the main priority. A model with native image or audio generation is a better fit when the required output is a media file rather than text. A different hosted model or service should be considered for web-search workflows, batch processing, fine-tuning or region-independent function calling, because those capabilities are not listed for this endpoint.

In short, Qwen3-VL-235B-A22B-Instruct is positioned as a high-capacity visual analysis model. Its appeal comes from the combination of difficult-document handling, video understanding, visual reasoning and open-weight access. Its main costs are the infrastructure required to run it, region-specific hosted features and the fact that its output remains text rather than generated media.


Answers to Frequently Asked Questions

How much does the Qwen3-VL-235B-A22B-Instruct API cost?
Alibaba Cloud lists US Virginia, Germany and China Beijing pricing at $0.287 per 1 million input tokens and $1.147 per 1 million output tokens. Singapore is listed at $0.400 per 1 million input tokens and $1.600 per 1 million output tokens. Actual charges may vary by region, promotions and token usage.
Can Qwen3-VL-235B-A22B-Instruct be self-hosted?
Yes. The model is available under the Apache 2.0 license on Hugging Face, with deployment guidance for Transformers, vLLM and SGLang. However, its 235 billion total parameters make self-hosting a demanding multi-GPU workload. Quantized and FP8 versions can reduce memory requirements but do not make it suitable for most ordinary consumer hardware.
What is the context window and maximum output length of Qwen3-VL-235B-A22B-Instruct?
The Alibaba Cloud Model Studio endpoint lists a 131,072-token context window, with a maximum input length of 129,024 tokens and a maximum output length of 32,768 tokens. Open-weight deployments describe a native 256K-token configuration and possible extensions up to 1 million tokens with YaRN, depending on hardware and serving software.
What are the main use cases for Qwen3-VL-235B-A22B-Instruct?
The model is suited to extracting information from difficult documents, analyzing charts and diagrams, reviewing videos, answering questions about images, converting screenshots into web code and supporting visual-agent workflows. It is most appropriate when visual accuracy and broad context matter more than minimal latency or low infrastructure cost.
What is Qwen3-VL-235B-A22B-Instruct?
Qwen3-VL-235B-A22B-Instruct is a large open-weight vision-language model from Alibaba Cloud's Qwen team. It processes text, images and video to perform tasks such as OCR, document analysis, visual question answering, video understanding, spatial reasoning and visual coding. Its primary output is text.


Sources 5
Provider

About Qwen