Qwen3.5-OCR

Qwen3.5-OCR

by Qwen · Current and available through Alibaba Cloud Model Studio

Qwen3.5-OCR is a specialized Alibaba Cloud Model Studio model for extracting text and information from images. It supports document parsing, text localization, table extraction, handwriting recognition, and key information extraction, with a 65,536-token context window and a maximum output of 16,384 tokens. China (Beijing) pricing is $0.069 per 1 million input tokens and $0.275 per 1 million output tokens. It is best suited to OCR pipelines rather than general chat, coding, tool use, web search, structured outputs, or media generation.

Text Reasoning Coding
Qwen3.5-OCR is designed for turning visual documents into usable text and structured information. Available through Alibaba Cloud Model Studio, it focuses on optical character recognition and document understanding: locating text in images, extracting content from tables, recognizing handwriting, and identifying important fields. Its documented context window is 65,536 tokens, with a maximum input length of 49,152 tokens and a maximum output length of 16,384 tokens. China (Beijing) pricing is listed at $0.069 per 1 million input tokens and $0.275 per 1 million output tokens. The model accepts image input and produces text output, but it is not intended to replace a general-purpose conversational, coding, tool-using, or multimodal generation model.
Outputs

What Qwen3.5-OCR can produce

Text
Inputs

What it can understand

Images
Model profile

Performance characteristics

4/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.5-OCR
Model type Multimodal
Context window 66K tokens
Maximum output 16K tokens
Status Current and available through Alibaba Cloud Model Studio
Model notes

Qwen3.5-OCR is a vision OCR model optimized for document parsing, text localization, and key information extraction. The official model page lists image input and text output only. It does not support function calling, structured outputs, web search, prefix completion, context caching, batch inference, or fine-tuning. The documented context window is 65,536 tokens, with maximum input length of 49,152 tokens and maximum output length of 16,384 tokens. Pricing shown is the original China (Beijing) API pricing and excludes limited-time promotions. Comparative scores are editorial estimates for this specialized OCR model, not provider benchmarks.

Cost

Model pricing

Input 0.069 USD per 1 million tokens in China (Beijing)
Output 0.275 USD per 1 million tokens in China (Beijing)
Model guide

Qwen3.5-OCR: Alibaba Cloud’s Specialized Model for Document Text Extraction

Qwen3.5-OCR is a specialized vision OCR model from Alibaba Cloud Model Studio. It accepts images and returns text, with capabilities focused on document parsing, text localization, table extraction, handwritten-text recognition, and key information extraction rather than general chat or coding.

What is Qwen3.5-OCR?

Qwen3.5-OCR is an OCR and visual document-understanding model provided by Alibaba Cloud through Model Studio. OCR, or optical character recognition, converts text appearing in an image into machine-readable text. Qwen3.5-OCR goes beyond simple character transcription by targeting related document tasks such as text localization, table extraction, handwritten-text recognition, and key information extraction.

For example, it can be used to process a scanned form, identify the words on a receipt, extract rows and columns from a table image, or recognize handwritten content. The model's role is therefore narrower and more operational than that of a general-purpose language model: its primary job is to read and interpret text contained in visual documents.

The supplied documentation identifies Qwen3.5-OCR as current and available through Alibaba Cloud Model Studio. No release date is provided in the available research, so its launch date cannot be verified here.

Core capabilities and supported modalities

Qwen3.5-OCR accepts image input and returns text output. The official model information supplied for this page lists no audio or video input and no image, audio, or video generation. In practical terms, it is a vision-to-text system rather than a model for creating media.

  • Image input: supported.
  • Text output: supported.
  • Audio input or output: not supported according to the supplied specifications.
  • Video input or output: not supported according to the supplied specifications.
  • Image generation: not supported.

Its documented uses include ordinary OCR, document parsing, text localization, table extraction, handwritten-text recognition, and key information extraction. Text localization is useful when an application needs to know not only what words appear in an image but also where those words are located. Key information extraction is useful for workflows that need to identify important values from forms or business documents, although the supplied research does not claim a separate structured-output interface.

Context, input, and output limits

Qwen3.5-OCR has a documented context window of 65,536 tokens. The supplied notes distinguish this overall context window from a maximum input length of 49,152 tokens and a maximum output length of 16,384 tokens.

SpecificationDocumented value
ProviderAlibaba Cloud
AvailabilityAlibaba Cloud Model Studio
Context window65,536 tokens
Maximum input length49,152 tokens
Maximum output length16,384 tokens
Primary inputImages
Primary outputText

These limits matter most when processing long or information-dense documents. A large document image or a multi-page workflow may require careful segmentation so that the request stays within the documented input limit. The maximum output length is generous for OCR and extraction tasks, but it does not imply that every request will produce a response of that size.

Where Qwen3.5-OCR is strongest

The model's main strength is specialization. Instead of treating OCR as a minor feature inside a broad chat model, Qwen3.5-OCR is positioned specifically for reading visual documents and extracting useful text from them.

  • Scanned documents: convert pages that do not contain selectable digital text into machine-readable content.
  • Receipts and forms: identify text and important fields from common business or administrative documents.
  • Tables: extract text from rows and columns in table images, subject to the quality and layout of the source.
  • Handwritten text: recognize handwriting where a conventional OCR system may be less suitable.
  • Text localization: identify text within an image and preserve information about its visual placement.
  • Document pipelines: provide text that can be passed to later search, classification, or data-processing stages.

These are provider-documented or research-supported use areas, not a claim that every document type will be read perfectly. Image quality, handwriting style, page layout, language, occlusion, and other factors can affect extraction accuracy. The supplied research does not provide benchmark scores, so no numerical accuracy claim should be inferred.

Pricing and cost considerations

The supplied Alibaba Cloud pricing is listed for the China (Beijing) region at $0.069 per 1 million input tokens and $0.275 per 1 million output tokens. These figures are the original China (Beijing) API prices and exclude limited-time promotions.

Token-based pricing can make the model attractive for high-volume OCR and extraction workloads, especially when requests are focused on converting images into concise text. Actual cost depends on the amount of image-related input and the length of the generated extraction. The available research does not provide a separate image-unit price, subscription plan, or regional price list, so those details should not be assumed from the token prices above.

The editorial cost score supplied for this model is 8 out of 10, while the editorial speed score is 7 out of 10. These are comparative estimates for this specialized OCR model, not Alibaba Cloud benchmark results or provider-published ratings. They suggest a favorable cost and speed trade-off for its intended document-processing role, but they should not be treated as a measured service-level guarantee.

Reasoning, coding, and tool support

Qwen3.5-OCR should not be selected as a general reasoning or programming model. The supplied editorial reasoning score is 4 out of 10 and the coding score is 1 out of 10. These scores are subjective evaluations of the model's suitability for those tasks, not official benchmark results.

Its useful reasoning is task-specific: interpreting a document, locating text, extracting fields, and organizing information found in an image. That is different from the broad, multi-step reasoning expected from a general-purpose language model. Likewise, while its text output may be consumed by software, the model is not documented as a coding assistant.

The supplied specifications list no function calling or tool use. They also list no structured-output capability, web search, prefix completion, context caching, batch inference, or fine-tuning. This makes Qwen3.5-OCR a poor fit for an application that expects the model itself to call external services, return provider-enforced JSON schemas, search the web, or participate in a large orchestration workflow. A separate application layer may still process the returned text, but that should not be confused with native model support.

Main limitations

Qwen3.5-OCR's specialization is also its main limitation. It is designed for image-based OCR and extraction, not for broad interactive assistance. The supplied research specifically identifies general-purpose chat, coding, function calling, web search, structured-output workflows, audio and video processing, and image generation as unsuitable uses.

There are also practical limitations to keep in mind:

  • It produces text rather than generated images, audio, or video.
  • It does not provide documented native tools or function calls.
  • It does not provide documented structured outputs, so applications requiring guaranteed schema-valid JSON may need additional validation or a different model.
  • It has no documented fine-tuning support for adapting the model to a specialized organization or document set.
  • It has no documented context caching or batch inference support in the supplied specifications.
  • The available research does not provide accuracy benchmarks or a knowledge cutoff.

For production use, developers should test representative documents rather than relying only on the model's task description. In particular, tables, handwriting, unusual layouts, low-resolution scans, and documents with overlapping visual elements deserve direct evaluation.

When to choose Qwen3.5-OCR

Choose Qwen3.5-OCR when the central problem is reading information from images and converting it into text. It is a sensible candidate for document ingestion, receipt processing, scanned archives, form extraction, table reading, and handwriting recognition. Its relatively low listed token prices and specialized purpose may be more appropriate than using a broad, expensive multimodal chat model for a high-volume OCR pipeline.

It is less appropriate when the main task is open-ended conversation, software development, web research, audio or video understanding, image generation, or tool-driven automation. In those cases, a general-purpose model or a model with native tool calling and structured-output support may be a better fit. Similarly, applications that require provider-guaranteed JSON schemas should not assume that Qwen3.5-OCR supplies that capability, because the available specifications list structured output as unsupported.

Overall, Qwen3.5-OCR is best understood as a focused visual document reader. Its value comes from concentrating on OCR and extraction tasks, with a 65,536-token context window and clear image-to-text positioning, rather than from offering the broad feature set of a general-purpose AI assistant.


Answers to Frequently Asked Questions

What is Qwen3.5-OCR used for?
Qwen3.5-OCR is an Alibaba Cloud model for extracting and interpreting text from images. It supports OCR, document parsing, text localization, table extraction, handwritten-text recognition, and key information extraction.
What input and output formats does Qwen3.5-OCR support?
Qwen3.5-OCR accepts images as input and returns text as output. According to the supplied specifications, it does not support audio or video input and output, image generation, or other media-generation tasks.
What are Qwen3.5-OCR's context, input, and output limits?
Qwen3.5-OCR has a documented context window of 65,536 tokens, with a maximum input length of 49,152 tokens and a maximum output length of 16,384 tokens.
How much does Qwen3.5-OCR cost?
The listed Alibaba Cloud API prices for the China (Beijing) region are $0.069 per 1 million input tokens and $0.275 per 1 million output tokens, excluding limited-time promotions. The available information does not provide separate image-unit pricing or prices for other regions.
What are the main limitations of Qwen3.5-OCR?
Qwen3.5-OCR is specialized for image-based OCR and document extraction rather than general-purpose chat, coding, web research, or tool-driven automation. The supplied specifications list no native function calling, web search, structured output, fine-tuning, context caching, or batch inference, and no accuracy benchmarks are provided.


Sources 2
Provider

About Qwen