Qwen3.5-Omni

Qwen3.5-Omni-Plus-Realtime

by Qwen · Current and available; canonical model identifier functionally equivalent to snapshot qwen3.5-omni-plus-realtime-2026-03-15

Alibaba Cloud's Qwen3.5-Omni-Plus-Realtime is a streaming multimodal model for voice assistants and visual conversations. It accepts text, images, video, and audio, returns text and speech, and supports controllable voices, interruption handling, function calling, web search, and real-time interfaces. Regional token pricing and limitations such as no structured outputs, batch API, fine-tuning, or context caching shape its best use cases.

Text Speech Reasoning Coding
Qwen3.5-Omni-Plus-Realtime is built for applications that need to listen, see, respond, and speak while an interaction is still taking place. Provided through Alibaba Cloud Model Studio, it handles text, images, video, and audio input and returns both text and generated speech. That makes it a better fit for live voice assistants, conversational customer service, and interactive multimedia analysis than for batch text generation or structured data extraction.
Outputs

What Qwen3.5-Omni-Plus-Realtime can produce

Text Speech
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
5/10 Coding
9/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.5-Omni
Model type Multimodal
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-03-15
Status Current and available; canonical model identifier functionally equivalent to snapshot qwen3.5-omni-plus-realtime-2026-03-15
Knowledge cutoff notes

Alibaba Cloud's model documentation does not provide a verified knowledge-cutoff date for this exact model.

Model notes

The exact API model ID is qwen3.5-omni-plus-realtime. It accepts text, images, video, and audio and returns text and audio. The model supports function calling, first-party web search, semantic interruption handling, controllable voice output, and custom voice cloning. Web search and function calling are mutually exclusive modes. Real-time access is documented through WebSocket, WebRTC, and AOQ interfaces. Alibaba Cloud lists 113 supported speech-recognition languages and dialects and 36 speech-generation languages and dialects. The canonical model is functionally equivalent to qwen3.5-omni-plus-realtime-2026-03-15, a snapshot dated March 15, 2026. Editorial capability scores are comparative estimates, not vendor-provided ratings.

Cost

Model pricing

Input China (Beijing): $1.38 per 1M text/image/video tokens; $11 per 1M audio tokens. Singapore: $2.10 per 1M text/image/video tokens; $16.50 per 1M audio tokens.
Output China (Beijing): $8.25 per 1M text-output tokens; $41.26 per 1M text-and-audio-output tokens. Singapore: $12.40 per 1M text-output tokens; $62 per 1M text-and-audio-output tokens.
Model guide

Qwen3.5-Omni-Plus-Realtime: Real-Time Voice and Visual Interaction

Qwen3.5-Omni-Plus-Realtime is Alibaba Cloud Model Studio's real-time multimodal model for streaming voice and visual applications. It accepts text, images, video, and audio, and produces text and audio responses through real-time interfaces. Its main differentiators are speech-to-speech interaction, controllable voice output, semantic interruption handling, function calling, web search, and support for long multimodal conversations.

What is Qwen3.5-Omni-Plus-Realtime?

Qwen3.5-Omni-Plus-Realtime is a real-time multimodal model from Alibaba Cloud Model Studio. Its canonical API identifier is qwen3.5-omni-plus-realtime. The model is designed for streaming interactions in which the system can receive spoken or visual context and respond without requiring separate speech-recognition and text-to-speech models.

In practical terms, an application can use it to hold a spoken conversation, interpret an image or video during that conversation, and produce a spoken answer. It can also return text for applications that need transcripts, captions, logs, or a conventional text interface alongside audio.

Alibaba Cloud lists the current model as functionally equivalent to the dated snapshot identifier qwen3.5-omni-plus-realtime-2026-03-15. The current model was released on March 15, 2026, according to the supplied Model Studio research.

Inputs, outputs, and voice capabilities

The model accepts four input types:

  • Text
  • Images
  • Video
  • Audio

Its documented outputs are text and audio. Audio output enables direct speech generation, while text output remains useful for transcripts, user-interface display, downstream processing, and application logging. Image and video generation are not supported by this model.

Alibaba Cloud documents speech recognition for 113 languages and dialects and speech generation for 36 languages and dialects. These are provider-published language-support figures; actual quality and voice availability may vary by language, deployment, and use case.

Voice output can be controlled with instructions affecting volume, speaking rate, and emotional delivery. The model also supports custom voice cloning, which may be useful for branded assistants or applications that require a particular voice identity. Voice cloning introduces additional consent, identity, and data-governance responsibilities, so it should not be treated as a purely technical feature.

Real-time interaction, interruption, and tools

Qwen3.5-Omni-Plus-Realtime is intended for streaming rather than one-shot requests. Alibaba Cloud documents real-time access through WebSocket, WebRTC, and AOQ interfaces, depending on the deployment and regional documentation. These interfaces are relevant when an application needs an ongoing conversation instead of repeatedly uploading complete recordings.

Semantic interruption handling is designed for conversations where a user starts speaking before the assistant has finished. This matters in voice interfaces because a usable assistant needs to recognize that a new utterance changes the conversation, rather than continuing to play a stale response.

The model supports function calling, allowing an application to expose defined operations such as checking an order, retrieving account information, or controlling an external system. It also supports Alibaba Cloud's first-party web search capability. The provider documents web search and function calling as mutually exclusive operating modes, so an implementation should choose the appropriate mode for each interaction rather than assuming both can be active simultaneously.

Context window and technical specifications

SpecificationDocumented value
Model identifierqwen3.5-omni-plus-realtime
Context window262,144 tokens
Maximum input length196,608 tokens
Maximum output length65,536 tokens
Input modalitiesText, image, video, and audio
Output modalitiesText and audio
Function callingSupported
Web searchSupported through a provider mode
Structured outputsNot supported
Context cachingNot supported
Batch APINot supported
Fine-tuningNot supported

The 262,144-token context window is substantially larger than the maximum input length because the full context budget also includes the model's possible output. The documented maximum output is 65,536 tokens, although a live voice application will normally produce much shorter responses to preserve conversational speed.

Structured outputs are not supported. Developers who need guaranteed JSON conforming to a schema should choose a model and API mode that explicitly provide structured-output support, or add application-side validation and recovery logic. The absence of context caching also makes this model less suitable for workflows that repeatedly reuse a large fixed prompt.

Pricing by region

Alibaba Cloud lists different token prices for China (Beijing) and Singapore. Pricing is separated by modality and output type rather than presented as one blended per-request rate.

RegionText, image, and video inputAudio inputText outputText and audio output
China (Beijing)$1.38 per 1 million tokens$11 per 1 million tokens$8.25 per 1 million tokens$41.26 per 1 million tokens
Singapore$2.10 per 1 million tokens$16.50 per 1 million tokens$12.40 per 1 million tokens$62 per 1 million tokens

Audio input and text-and-audio output are priced above ordinary text, image, and video input. This means that a speech-heavy application can have a very different cost profile from an application that mainly sends text or visual context and receives text responses. The applicable region, billing rules, minimums, and any account-level limits should be confirmed in the current Alibaba Cloud Model Studio pricing documentation before deployment.

Capability, speed, and cost trade-offs

The supplied comparative assessment gives Qwen3.5-Omni-Plus-Realtime a speed score of 9 out of 10, a reasoning score of 6, a coding score of 5, and a cost score of 5. These are editorial estimates, not Alibaba Cloud benchmarks or provider-published ratings. They are best understood as a summary of the model's intended trade-off: fast interactive behavior and broad multimodal handling are more central than maximum coding or reasoning performance at the lowest possible price.

For a live assistant, the ability to process audio and respond quickly may matter more than achieving the strongest possible performance on difficult text-only reasoning tasks. Conversely, an application that only needs asynchronous document processing may pay for real-time audio features it never uses. The model's output pricing also makes spoken responses more expensive than text-only responses, particularly in Singapore.

Reasoning, coding, and important limitations

Qwen3.5-Omni-Plus-Realtime can support reasoning within multimodal conversations, such as interpreting a visual scene, following a spoken instruction, or deciding when to invoke a defined function. However, the supplied research does not provide a standardized reasoning benchmark for this exact model. Its editorial reasoning rating should not be read as a verified benchmark result.

Coding is supported as a general capability of the model ecosystem, but this real-time model is not primarily positioned as a code-generation specialist. Its documented limitations include the lack of structured outputs, batch inference, fine-tuning, and persistent prompt caching. It is also not an image-generation, video-generation, embedding, or music-generation model.

Function calling does not make the model's actions automatically safe or correct. The surrounding application must validate arguments, enforce permissions, handle failed calls, and decide whether a spoken instruction is sufficiently clear to trigger an external action. Similarly, web search can provide current information, but retrieved content and generated answers still require application-appropriate verification.

Best use cases

This model is a strong fit when an application needs several of the following at the same time:

  • Voice assistants that listen and speak in real time.
  • Customer-service agents that need interruption handling and tool access.
  • Visual conversational agents that can discuss images or video while speaking with a user.
  • Live multimedia analysis, such as asking questions about an incoming recording or visual scene.
  • Spoken interfaces with controllable volume, pace, or emotional delivery.
  • Applications that need text transcripts alongside generated speech.
  • Real-time assistants that can use either function calling or provider web search.

For example, a support assistant could receive a customer's spoken explanation, inspect an uploaded product image, call an order-status function, and answer with both a displayed transcript and spoken guidance. The application would still need to manage authentication, tool permissions, conversation state, and response validation.

When to choose Qwen3.5-Omni-Plus-Realtime

Choose Qwen3.5-Omni-Plus-Realtime when low-latency multimodal conversation is the central requirement. It is particularly appropriate when audio input and audio output are first-class parts of the experience, and when the ability to interpret visual context or call application tools adds value.

Another model type may be more appropriate when the workload is mainly text-only reasoning, large-scale batch processing, strict JSON generation, embeddings, fine-tuning, or persistent prompt reuse. A text-focused model can avoid the cost and complexity of audio streaming when speech is not needed. A model with explicit structured-output support is preferable for workflows where downstream software requires schema-valid JSON. For image or video generation, a dedicated generative model is required because Qwen3.5-Omni-Plus-Realtime only accepts those media types as input.

The main practical choice is therefore not simply whether the model supports many modalities. It is whether the application benefits enough from live speech, visual context, and interruption-aware interaction to justify the model's regional pricing and the additional engineering required by real-time APIs.


Answers to Frequently Asked Questions

How much does Qwen3.5-Omni-Plus-Realtime cost?
Pricing varies by region and modality. In China (Beijing), text, image, and video input costs $1.38 per 1 million tokens, audio input costs $11, text output costs $8.25, and text-and-audio output costs $41.26. In Singapore, the corresponding prices are $2.10, $16.50, $12.40, and $62 per 1 million tokens. Current Alibaba Cloud pricing and account-level billing rules should be confirmed before deployment.
Can Qwen3.5-Omni-Plus-Realtime use function calling and web search?
Yes, the model supports function calling and Alibaba Cloud's first-party web search capability. However, the provider documents these as mutually exclusive operating modes, so an application must choose the appropriate mode for each interaction.
Does Qwen3.5-Omni-Plus-Realtime support real-time voice conversations and interruptions?
Yes. The model is intended for streaming voice interactions and supports real-time access through WebSocket, WebRTC, and AOQ interfaces, depending on the deployment. It also provides semantic interruption handling so users can begin speaking before the assistant finishes responding.
What is Qwen3.5-Omni-Plus-Realtime used for?
Qwen3.5-Omni-Plus-Realtime is designed for real-time multimodal applications such as voice assistants, customer-service agents, visual conversational systems, live multimedia analysis, and spoken interfaces that can use tools. It can process text, images, video, and audio while returning text and audio responses.
What input and output modalities does Qwen3.5-Omni-Plus-Realtime support?
The model accepts text, images, video, and audio as input. Its documented outputs are text and audio. It does not support image generation or video generation.


Sources 5
Provider

About Qwen