Qwen3.5-Omni

Qwen3.5-Omni-Plus

by Qwen · current

Qwen3.5-Omni-Plus is Alibaba Cloud's multimodal model for understanding text, images, audio and video while producing text and natural speech. It supports long audio and video inputs, streaming, web search and region-dependent function calling and batch inference. Its main trade-offs are modality-dependent pricing, inconsistent regional tool support, unsupported structured outputs and no image or video generation.

Text Speech Reasoning Coding
Qwen3.5-Omni-Plus is an end-to-end multimodal model available through Alibaba Cloud Model Studio. It accepts text, images, audio and video, then returns text, synthesized speech, or both. Its main distinction is that speech generation is part of the model's output rather than an entirely separate text-to-speech step. The model also supports long audio and video inputs, streaming responses, web search and selected tool features, although availability varies by deployment region.
Outputs

What Qwen3.5-Omni-Plus can produce

Text Speech
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Batch API Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
6/10 Coding
7/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.5-Omni
Model type Multimodal
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-03-15
Status current
Knowledge cutoff notes

Alibaba Cloud's public model documentation reviewed for this record does not state a verified knowledge cutoff for Qwen3.5-Omni-Plus.

Model notes

The canonical model ID qwen3.5-omni-plus is currently functionally equivalent to qwen3.5-omni-plus-2026-03-15. It accepts text, images, audio, and video and returns text and speech audio. Documentation describes up to three hours of audio input, up to one hour of video input, and files up to 2 GB in the Qwen3.5-Omni API workflow. Function calling is supported for the China (Beijing) Chat Completions deployment with text output but is not listed for the international scope. Web search is supported. Batch inference is supported in China (Beijing) but not in the Singapore/international scope listing. Structured outputs and context caching are listed as unsupported on the exact model page. Audio output is billed using modality-specific token accounting, and text in text-and-audio output is not charged separately on the international pricing table. Editorial scores are comparative estimates, not provider-published ratings.

Cost

Model pricing

Input $1.40 per 1M tokens for text/image/video input; $11.00 per 1M tokens for audio input in international deployments. China (Beijing) snapshot pricing is $0.96 per 1M tokens for text/image/video input and $7.29 per 1M tokens for audio input.
Output $8.30 per 1M tokens for text output; $44.00 per 1M tokens for text-and-audio output in international deployments. China (Beijing) snapshot pricing is $5.50 per 1M tokens for text output and $29.29 per 1M tokens for text-and-audio output.
Model guide

Qwen3.5-Omni-Plus: A Multimodal Model for Speech, Audio and Video

Qwen3.5-Omni-Plus is Alibaba Cloud's multimodal Qwen model for understanding text, images, audio and video while producing text and natural speech. It is aimed at voice assistants, multimedia analysis, accessibility tools and other applications that need audio-visual understanding with spoken responses.

What is Qwen3.5-Omni-Plus?

Qwen3.5-Omni-Plus is a multimodal model from Alibaba's Qwen3.5-Omni family, provided through Alibaba Cloud Model Studio. Multimodal means that the model can work with more than text: a request can include written instructions together with images, audio recordings, video or combinations of these formats.

The model is particularly focused on audio-visual interaction. It can analyze spoken content and video, answer questions about media, and produce a spoken response as well as text. That makes it more suitable for voice assistants and media-analysis applications than a text-only model, while its speech output distinguishes it from multimodal models that only describe audio or video in writing.

The current rolling model identifier is qwen3.5-omni-plus. Alibaba Cloud documentation states that it is functionally equivalent to the dated snapshot qwen3.5-omni-plus-2026-03-15. The dated snapshot documentation specifies a 262,144-token context window, with a maximum input length of 196,608 tokens and a maximum output length of 65,536 tokens.

Supported input and output modalities

Qwen3.5-Omni-Plus accepts text, images, audio and video. These inputs can be combined, allowing an application to ask a question about an image, provide an audio explanation alongside a video, or submit a media file with written instructions.

  • Text input: written prompts and instructions.
  • Image input: image understanding within multimodal requests.
  • Audio input: speech and other audio analysis, including long recordings.
  • Video input: video understanding and audio-visual analysis.
  • Text output: written answers and analysis.
  • Audio output: synthesized speech for spoken responses.

Alibaba Cloud's documentation describes audio input of up to three hours and video input of up to one hour, with files up to 2 GB in the documented Qwen3.5-Omni API workflow. These are workflow and documentation limits, so an implementation should verify the selected endpoint and request method before assuming that every deployment exposes identical limits.

The provider describes audio input support in more than 60 languages and speech output in more than 30 languages. The exact language and dialect coverage should be checked against the current model documentation when building a multilingual product.

Where it fits in the Qwen lineup

Qwen3.5-Omni-Plus is a specialized member of the broader Qwen model family rather than a general-purpose image or video generation model. Its role is understanding mixed media and responding with text or speech. Although Alibaba's wider Qwen ecosystem includes models and services for other tasks, this model's defining capability is the combination of audio-visual input with native speech output.

This positioning matters when selecting a model. Qwen3.5-Omni-Plus can explain a video aloud, summarize a recording, or support a spoken interaction. It should not be selected when the central requirement is generating a new image or video, because its documented non-text output is speech audio rather than generated images or video.

Context, API access and tools

Qwen3.5-Omni-Plus is available through Alibaba Cloud Model Studio using Chat Completions. Streaming is supported, including streamed audio response data. Streaming can be useful for conversational applications because the client can begin receiving a response before the complete generation has finished.

The documented 262,144-token context window provides room for long prompts and substantial media-related context. However, the input and output limits are separate: the dated model documentation specifies up to 196,608 input tokens and up to 65,536 output tokens. Multimedia files also have modality-specific processing and billing rules, so token capacity should not be treated as a simple measure of file duration or size.

Web search is listed as supported. Function calling, which allows a model to request actions from external software, is more limited: Alibaba Cloud lists it for the China (Beijing) deployment with text output, but the international scope documentation does not list it as supported. Batch inference is listed as supported in China (Beijing) and not in the Singapore or international scope listing.

Structured outputs and context caching are listed as unsupported on the exact model page. Applications that require strict JSON Schema enforcement or provider-managed prompt caching should therefore verify the endpoint carefully and consider whether another model or an application-side validation layer is more appropriate.

Qwen3.5-Omni-Plus pricing

Alibaba Cloud bills the model by tokens, with separate rates for different input and output modalities. The current international pricing listed in the supplied documentation is:

Usage typeInternational price
Text, image or video input$1.40 per 1 million tokens
Audio input$11.00 per 1 million tokens
Text output$8.30 per 1 million tokens
Text and audio output$44.00 per 1 million tokens

For text-and-audio output, the international pricing information states that the text portion is not charged separately. Multimedia tokenization means that the apparent cost of a request depends on the type and amount of media being processed, not only on the visible length of the prompt.

The dated snapshot documentation lists lower China (Beijing) rates: $0.96 per million tokens for text, image or video input, $7.29 per million tokens for audio input, $5.50 per million tokens for text output and $29.29 per million tokens for text-and-audio output. Prices and regional availability can change, so production deployments should use the current Alibaba Cloud pricing page for final estimates.

Main strengths and trade-offs

The most important strength is the combination of broad media understanding and spoken output. A developer can build an assistant that listens to a user, examines an image or video, and replies conversationally without treating speech as an entirely separate product feature. Long documented audio and video limits also make the model relevant to recordings, media review and extended spoken interactions.

Its large context allowance is another practical advantage for applications that need to combine lengthy instructions with substantial conversation or media-related context. Streaming can improve the user experience in voice interfaces, while web search can support tasks that need current information where the feature is available.

The trade-offs are cost, regional inconsistency and output specialization. Audio input is substantially more expensive than text, image or video input in the international pricing table, and text-plus-audio output costs more than text-only output. Function calling and batch inference are not uniformly available across deployments. The model also does not provide image or video generation and does not advertise structured output support on its model page.

The editorial assessment for this record rates reasoning at 7 out of 10, coding at 6 out of 10, speed at 7 out of 10 and cost at 5 out of 10. These are comparative editorial estimates, not Alibaba Cloud benchmark results or provider-published ratings. The model's strongest case is multimodal and speech interaction, rather than the lowest-cost text generation or strict software-engineering workflows.

Best use cases

  • Voice assistants: applications that need spoken, multilingual answers rather than text alone.
  • Audio-visual analysis: summarizing meetings, recordings, lectures or video content.
  • Spoken explanations: describing an image, document or video for hands-free use.
  • Accessibility tools: interfaces that convert visual or written information into conversational audio.
  • Customer support: systems that combine voice interaction with image, audio or video attachments.
  • Interactive media workflows: reviewing, monitoring or querying multimedia content.
  • Multilingual applications: products requiring audio understanding and speech responses across multiple languages.

For example, a media-review application could accept a video, identify important moments, return a written summary and stream a spoken explanation. A support assistant could receive a user's voice recording and an image of a device, then answer with both text instructions and audio guidance.

When to choose this model

Choose Qwen3.5-Omni-Plus when the application genuinely needs several input modalities and a natural audio response. It is a strong candidate for voice-first products, multimedia assistants, spoken accessibility features and applications that analyze long audio or video.

A different option may be more appropriate when the task is text-only and cost-sensitive, when strict structured JSON output is essential, or when function calling and batch processing must work consistently across international regions. A text-focused model can also be simpler and less expensive for ordinary chat, summarization or coding. Likewise, an image or video generation model is the better category for creating visual assets, because Qwen3.5-Omni-Plus is documented as an understanding and speech-output model, not a visual-generation model.

Before deployment, verify the target region, model identifier, current price, supported function-calling behavior, batch availability, audio language coverage and file limits. Those details are not uniform across all Alibaba Cloud Model Studio scopes.

Bottom line

Qwen3.5-Omni-Plus is best understood as a speech-enabled multimodal model: it can interpret text, images, audio and video, then respond in text and synthesized speech. Its long-media support, streaming and multilingual speech capabilities make it useful for voice and audio-visual applications. The principal compromises are higher modality-dependent costs, region-specific tools, unsupported structured outputs and the absence of image or video generation.


Answers to Frequently Asked Questions

What are the best use cases for Qwen3.5-Omni-Plus?
The model is well suited to voice assistants, meeting and video summarization, spoken explanations, accessibility tools, multimedia customer support, interactive media review and multilingual applications that need audio understanding and speech responses.
How much does Qwen3.5-Omni-Plus cost?
The listed international rates are $1.40 per million tokens for text, image or video input, $11.00 for audio input, $8.30 for text output and $44.00 for text-and-audio output. Pricing varies by region and may change, so production estimates should use the current Alibaba Cloud pricing page.
What input and output modalities does Qwen3.5-Omni-Plus support?
The model supports text, image, audio and video input. Its outputs include text and audio speech, but it does not generate new images or videos.
How much audio and video can Qwen3.5-Omni-Plus process?
Alibaba Cloud documentation describes support for audio recordings up to three hours and videos up to one hour, with files up to 2 GB in the documented Qwen3.5-Omni API workflow. Limits may vary by endpoint and request method.
What is Qwen3.5-Omni-Plus?
Qwen3.5-Omni-Plus is a multimodal Alibaba Cloud model that accepts text, images, audio and video, then responds with written answers or synthesized speech. It is designed for audio-visual analysis, voice assistants and spoken interactions.


Sources 4
Provider

About Qwen