Qwen3.5-Omni

Qwen3.5-Omni-Flash-Realtime

by Qwen · Current and available; rolling model identity functionally equivalent to qwen3.5-omni-flash-realtime-2026-03-15

A practical guide to Qwen3.5-Omni-Flash-Realtime, covering its WebSocket realtime design, multimodal inputs, text and speech outputs, context and session limits, tool support, pricing, strengths, limitations, and best use cases.

Text Speech Reasoning Coding
Qwen3.5-Omni-Flash-Realtime is a WebSocket-based real-time model from Alibaba Cloud Model Studio. It accepts text, images, video, and streaming audio, then responds with text, spoken audio, or both. The model is aimed at applications where conversational delay matters, including voice assistants, customer-service agents, tutoring systems, and multimodal interactive agents. It combines fast streaming behavior with a large context window, but it is not the best fit for fine-tuning, batch workloads, schema-constrained JSON, or applications centered on deep reasoning and specialized code generation.
Outputs

What Qwen3.5-Omni-Flash-Realtime can produce

Text Speech
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.5-Omni
Model type Multimodal
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-03-26
Status Current and available; rolling model identity functionally equivalent to qwen3.5-omni-flash-realtime-2026-03-15
Knowledge cutoff notes

Alibaba Cloud's public model documentation does not provide a verified knowledge cutoff date for this exact realtime model. Its web search capability can retrieve current information during use, but that does not establish or change the underlying model knowledge cutoff.

Model notes

The model uses the Realtime API over WebSocket and supports text, image, video, and streaming audio input with text and audio output. Alibaba Cloud documents support for function calling, web search, controllable voice dialogue, semantic interruption, and voice cloning. Web search and function calling cannot be enabled simultaneously in one session. The current rolling model ID is functionally equivalent to the dated snapshot qwen3.5-omni-flash-realtime-2026-03-15. The documented context window is 262,144 tokens, with a 196,608-token maximum input length and 65,536-token maximum output length. Retained realtime history is limited to 80 audio turns and 50 video turns, and a WebSocket session can last up to 120 minutes. Audio and visual inputs are converted into token usage for billing. Editorial scores are comparative estimates rather than vendor benchmarks.

Cost

Model pricing

Input International: $0.55 per 1M tokens for text/image/video input; $4.50 per 1M tokens for audio input. China mainland: $0.45 per 1M tokens for text/image/video input; $3.71 per 1M tokens for audio input.
Output International: $3.30 per 1M tokens for text output; $17.70 per 1M tokens for audio output. China mainland: $2.75 per 1M tokens for text output; $14.71 per 1M tokens for audio output.
Model guide

Qwen3.5-Omni-Flash-Realtime: Built for Low-Latency Multimodal Voice Interaction

Qwen3.5-Omni-Flash-Realtime is Alibaba Cloud Model Studio's real-time multimodal model for streaming text, audio, images, and video into interactive sessions while producing text or natural speech audio. Its WebSocket Realtime API, voice activity detection, semantic interruption handling, web search, and function calling make it particularly suitable for voice assistants and interactive multimedia applications, although its tool features have session-level constraints and its audio output is substantially more expensive than text output.

What Qwen3.5-Omni-Flash-Realtime is

Qwen3.5-Omni-Flash-Realtime is the lower-latency real-time member of Alibaba Cloud Model Studio's Qwen3.5-Omni family. Alibaba Cloud lists the canonical model ID as qwen3.5-omni-flash-realtime. The rolling model identity is documented as functionally equivalent to the dated snapshot qwen3.5-omni-flash-realtime-2026-03-15.

The model is designed for ongoing, interactive sessions rather than isolated text prompts. It connects through the Realtime API over WebSocket, a communication method that keeps a live connection open so audio and other events can be exchanged incrementally. This allows an application to stream microphone audio, receive partial responses, detect when a person has started speaking, and interrupt an answer when the conversation changes direction.

Alibaba Cloud's release catalog identifies the model as current and available, with a listed release or catalog date of March 26, 2026. The specifications and prices described here are based on the supplied Alibaba Cloud Model Studio documentation. Editorial assessments of speed, cost, reasoning, and coding are separate comparative judgments, not vendor benchmark scores.

Inputs, outputs, and supported modalities

Qwen3.5-Omni-Flash-Realtime can receive four broad types of input:

  • Text messages
  • Images
  • Video, handled through extracted image frames
  • Streaming audio

It can produce text and audio. That makes speech-to-speech interaction possible, but it can also return a text transcript or text-only answer when spoken output is unnecessary. The documented basic audio configuration uses 16 kHz, 16-bit, mono PCM for input and 24 kHz, 16-bit, mono PCM for output.

For video, the service processes frames rather than treating a video file as an unexplained continuous object. The documentation recommends approximately one frame per second for many use cases. The appropriate frame rate depends on the application: a slowly changing presentation may need fewer frames, while a rapidly changing scene may require more frequent sampling and will create more input usage.

The model documentation also lists controllable voice dialogue, semantic interruption handling, and voice cloning support within the Qwen3.5-Omni realtime service. These features are relevant to applications in which the assistant must behave like a live participant instead of waiting for a complete request before responding.

Why the realtime design matters

A conventional request-and-response API normally waits for an application to submit a complete request. Qwen3.5-Omni-Flash-Realtime is intended for a different interaction pattern. Audio can be sent as it arrives, and the model can stream back an answer while the user is still engaged in the conversation.

Semantic interruption is especially important for voice interfaces. Rather than forcing a user to wait for the assistant to finish speaking, an application can respond when the user begins a meaningful interruption. This can make customer-service systems, hands-free assistants, and tutoring applications feel more natural. Developers still need to manage the WebSocket lifecycle, audio buffers, response events, voice settings, and interruption state themselves.

The model supports both web search and function calling. Web search can help retrieve current information, while function calling allows the model to request an operation from an external application, such as checking an appointment or controlling a workflow. Alibaba Cloud's current realtime documentation states that web search and function calling are mutually exclusive within a session. An implementation therefore needs to choose which capability is active for a particular session rather than assuming both can be enabled simultaneously.

Context window and session limits

The model has a documented context window of 262,144 tokens. Alibaba Cloud lists a maximum input length of 196,608 tokens and a maximum output length of 65,536 tokens. A token is a unit of text or multimodal content used for processing and billing; it does not correspond exactly to a word.

These large limits can be useful for extended conversations or sessions that combine transcripts, images, and other context. They should not be interpreted as unlimited memory. The realtime documentation separately lists retained-history limits of up to 80 audio turns and 50 video turns. A single WebSocket session can last up to 120 minutes.

Audio and visual material are converted into token usage for billing. Long audio sessions, high-frequency video frames, and unnecessarily large image inputs can therefore increase both processing volume and cost. Applications should retain only the context needed for the current task and select a suitable frame rate for video.

Pricing and cost trade-offs

Alibaba Cloud publishes separate international and China mainland rates by modality. The current international rates are:

Usage typeInternational priceChina mainland price
Text, image, and video input$0.55 per 1 million tokens$0.45 per 1 million tokens
Audio input$4.50 per 1 million tokens$3.71 per 1 million tokens
Text output$3.30 per 1 million tokens$2.75 per 1 million tokens
Audio output$17.70 per 1 million tokens$14.71 per 1 million tokens

The most important practical cost distinction is between text and audio. International audio output is priced at $17.70 per million tokens, compared with $3.30 per million tokens for text output. An application that only needs a written answer should avoid generating speech. Conversely, the additional audio cost may be justified when the product's value depends on hands-free or spoken interaction.

Input prices also vary significantly by modality. Audio input costs more than text, image, or video input under the listed international rates. Since visual and audio data are converted into tokens, developers should test representative sessions rather than estimating cost from message count alone.

Reasoning, coding, and tool capabilities

Qwen3.5-Omni-Flash-Realtime is primarily a fast multimodal interaction model, not a specialist reasoning or coding model. The supplied editorial assessment gives it a reasoning score of 6 out of 10 and a coding score of 6 out of 10. These are comparative editorial estimates, not scores published by Alibaba Cloud and not standardized benchmark results.

It can support coding-related conversations and use external functions, but its strongest role is coordinating a live multimodal exchange. For example, it may be appropriate for a voice assistant that asks a backend service for information, or for an interactive agent that interprets a user's spoken request alongside an image. A workload requiring extensive mathematical reasoning, highly reliable code generation, or strict structured output may be better served by another model or API mode.

Structured outputs are not listed as supported for this exact realtime model. Fine-tuning, batch inference, and context caching are also not listed as supported. This means developers should not assume that a response can be constrained to a particular JSON schema simply because the model can call functions. Function calling and schema-constrained response generation are related but distinct capabilities.

Main strengths and limitations

Strengths

  • Low-latency interaction: Streaming audio and WebSocket sessions are designed for conversational applications where waiting for a complete response is undesirable.
  • Broad multimodal input: The model can combine text, images, video frames, and live audio in one interactive experience.
  • Speech output: It can return natural audio as well as text, supporting speech-to-speech applications.
  • Conversation control: Voice activity detection, semantic interruption, and controllable voice dialogue support more responsive assistants.
  • Tools and current information: Web search and function calling are available, subject to the restriction that they cannot be used together in one session.
  • Large documented context: The 262,144-token context window is substantial for long interactive sessions, subject to separate turn and duration limits.

Limitations

  • Realtime implementation complexity: Developers must handle WebSocket connections, streaming events, audio formats, buffering, and interruptions.
  • Tool exclusivity: A session must choose between web search and function calling according to the current documentation.
  • No listed structured output: The model is not documented as supporting schema-constrained JSON for this exact realtime endpoint.
  • Higher speech costs: Audio input and especially audio output cost more than text under the international pricing schedule.
  • Limited specialized workflows: Fine-tuning, batch inference, and context caching are not listed as supported.
  • Finite realtime sessions: History is limited to 80 audio turns and 50 video turns, and one WebSocket session can last up to 120 minutes.

Best use cases

Qwen3.5-Omni-Flash-Realtime is a strong candidate when the application needs immediate spoken interaction and multimodal understanding at the same time. Suitable uses include:

  • Voice assistants that can see images or video supplied by the user
  • Customer-service agents that respond verbally and invoke a backend function
  • Real-time tutoring or coaching with spoken questions and visual material
  • Multilingual voice interfaces and hands-free applications
  • Interactive media analysis, such as discussing a video or image during playback
  • Multimodal agents that need to react to interruptions and changing user intent

For a voice assistant, the model can stream a user's speech, interpret the request, optionally call an external service, and return spoken audio. For an image-based assistant, a user might provide an image while speaking a question about it. The model's value in these examples comes from combining modalities within a live conversation, not simply from generating a longer text answer.

When to choose this model

Choose Qwen3.5-Omni-Flash-Realtime when low conversational latency, speech output, and multimodal input are central product requirements. It is particularly attractive when a text-only model would make the interaction feel slow or inconvenient, and when the application can justify the additional cost and engineering work of realtime audio.

Choose a different option when the priority is strict JSON schema compliance, offline or batch processing, fine-tuning, extensive deep reasoning, or specialized code generation. A standard non-realtime model may also be more appropriate when users submit complete requests and do not benefit from live interruption or streaming speech. Likewise, if the application rarely needs audio output, text output will generally offer a lower-cost path under the published rates.

Within a larger system, Qwen3.5-Omni-Flash-Realtime can therefore be treated as the conversational front end rather than the universal model for every task. A product may use it for speech and multimodal interaction while routing demanding reasoning, structured data generation, or batch work to a more specialized model. That division is a practical trade-off, not a claim that the realtime model cannot discuss code or perform reasoning.

Bottom line

Qwen3.5-Omni-Flash-Realtime is designed for live multimodal conversation: it streams audio, understands text and visual inputs, and responds with text or speech through a WebSocket Realtime API. Its most distinctive features are low-latency voice interaction, semantic interruption, controllable dialogue, and support for web search or function calling. Its principal trade-offs are implementation complexity, modality-based billing, session limits, and the absence of documented structured outputs, fine-tuning, and batch support. For voice-first assistants and interactive multimedia agents, those trade-offs may be worthwhile; for conventional text generation or highly constrained data-processing workflows, another model type is likely to be a better fit.


Answers to Frequently Asked Questions

What are the main limitations of Qwen3.5-Omni-Flash-Realtime?
The model requires developers to manage WebSocket connections, streaming events, audio buffers, and interruptions. It does not list structured outputs, fine-tuning, batch inference, or context caching as supported for this endpoint. Sessions are limited to 120 minutes, with retained history of up to 80 audio turns and 50 video turns.
Can Qwen3.5-Omni-Flash-Realtime use web search and function calling together?
No. According to Alibaba Cloud's current realtime documentation, web search and function calling are mutually exclusive within a session. An application must choose which capability to enable for each session.
How much does Qwen3.5-Omni-Flash-Realtime cost?
For international usage, the listed prices are $0.55 per 1 million tokens for text, image, and video input; $4.50 for audio input; $3.30 for text output; and $17.70 for audio output. China mainland rates are listed separately and are lower. Actual costs depend on the amount and modality of content processed.
What is Qwen3.5-Omni-Flash-Realtime?
Qwen3.5-Omni-Flash-Realtime is Alibaba Cloud Model Studio's low-latency multimodal model for live voice interactions. It uses the Realtime API over WebSocket to stream audio, process text and visual inputs, detect interruptions, and return text or spoken responses.
What input and output modalities does Qwen3.5-Omni-Flash-Realtime support?
The model accepts text, images, video through extracted frames, and streaming audio. It can produce text and audio, enabling speech-to-speech applications as well as text-only responses.


Sources 5
Provider

About Qwen