Qwen3.8-Omni

Qwen3.8-Omni-Flash-Realtime

by Qwen · Current and available; international deployment in Singapore

A factual overview of Qwen3.8-Omni-Flash-Realtime, including its real-time audio and video capabilities, text and speech output, context and history limits, pricing, tool support, regional availability, strengths, limitations, and practical use cases.

Text Speech Reasoning Coding
Qwen3.8-Omni-Flash-Realtime is designed for real-time voice and video interaction rather than conventional request-response inference. It accepts text, streaming audio, and video represented as consecutive image frames, and can return text and spoken audio with a maximum output length of 65,536 tokens.
Outputs

What Qwen3.8-Omni-Flash-Realtime can produce

Text Speech
Inputs

What it can understand

Text Audio Video Multimodal input
Capabilities

Supported features

Tool use Streaming Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.8-Omni
Model type Multimodal
Context window 197K tokens
Maximum output 66K tokens
Release date 2026-09-21
Status Current and available; international deployment in Singapore
Knowledge cutoff notes

Alibaba Cloud's current model documentation does not publish a specific knowledge-cutoff date for this real-time model.

Model notes

The exact API model ID is qwen3.8-omni-flash-realtime. It accepts text, streaming audio, and video represented as consecutive image frames, and returns text and audio. The model supports custom function calling, remote MCP tools, multichannel audio through WebSocket, video aggregation, and WebSocket, WebRTC, or AOQ access. Audio history is limited to 100 turns and 600 seconds, while video history is limited to 50 turns and 240 seconds; older history is discarded when cumulative limits are exceeded. Speech output charges both the corresponding text and audio outputs. A 1-million-token free quota is available only in China (Beijing) for 90 days subject to Alibaba Cloud's eligibility rules. The official model page lists China (Beijing) and Singapore as supported regions. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input China (Beijing): CNY 1.5 per 1 million tokens for text/images/video input and CNY 6 per 1 million tokens for audio input. Singapore: CNY 1.677 per 1 million tokens for text/images/video input and CNY 6.781 per 1 million tokens for audio input.
Output China (Beijing): CNY 4.5 per 1 million tokens for text output and CNY 12 per 1 million tokens for audio output. Singapore: CNY 5.104 per 1 million tokens for text output and CNY 13.636 per 1 million tokens for audio output.
Model guide

Qwen3.8-Omni-Flash-Realtime: Real-Time Speech-to-Speech and Video Interaction

Qwen3.8-Omni-Flash-Realtime is Alibaba Cloud Model Studio's real-time multimodal model for interactive text, streaming audio, and video sessions. It produces both text and audio, supports multichannel audio, custom function calling, remote MCP tools, WebSocket, WebRTC, and AOQ access.

What is Qwen3.8-Omni-Flash-Realtime?

Qwen3.8-Omni-Flash-Realtime is Alibaba Cloud Model Studio's real-time multimodal model for applications that need to listen, interpret, and respond during an ongoing session. Its main purpose is interactive speech and video processing, not ordinary one-off text generation.

The model can receive text, streaming audio, and video represented as a sequence of image frames. It can respond with text and generated audio, making it suitable for speech-to-speech assistants, live video agents, meeting interfaces, and other applications where waiting for a conventional request-response cycle would make the experience feel slow.

The official model identifier is qwen3.8-omni-flash-realtime. The supplied documentation lists China (Beijing) and Singapore as supported regions, with the international deployment available in Singapore. The model was released on September 21, 2026, according to the supplied lifecycle documentation.

Where it fits in Alibaba Cloud's model catalog

Qwen3.8-Omni-Flash-Realtime belongs to the Qwen3.8-Omni family and is positioned in Alibaba Cloud Model Studio's speech-to-speech and omni-modal categories. That positioning matters: it is built around continuous audio and video interaction rather than being a general-purpose text model with occasional audio features.

Alibaba Cloud Model Studio provides the model through real-time interfaces including WebSocket, WebRTC, and AOQ. WebSocket is useful for applications that need a persistent bidirectional connection, while WebRTC is oriented toward low-latency browser or media communication. The documentation also lists multichannel audio support through WebSocket.

These are verified interface and catalog details from the supplied Alibaba Cloud documentation. They should not be interpreted as a guarantee that every interface has identical regional availability, limits, or SDK behavior.

Supported input and output modalities

The model is multimodal in a specific, real-time sense:

  • Text input: Supported.
  • Audio input: Supported as streaming audio.
  • Video input: Supported through consecutive image frames representing video.
  • Text output: Supported.
  • Audio output: Supported, including spoken responses.
  • Image or video generation: Not supported as an output capability.

The research identifies audio and video input, but does not list standalone still-image input as a separate model capability. Video is handled as a stream or sequence of frames rather than as a video-generation task. Likewise, the model can produce spoken audio but is not an image, video, music, or embedding generator.

In practical terms, an application could stream a user's voice while also supplying video frames, then receive both text for display and audio for playback. This makes the model more appropriate for a live assistant than for a workflow that only needs a written answer after a file upload.

Context, history, and output limits

Qwen3.8-Omni-Flash-Realtime has a documented context length of 196,608 tokens and a maximum output length of 65,536 tokens. A token is a unit of text or encoded media information used by the model, so these figures should not be treated as a direct character, word, minute, or file-size limit.

The real-time history limits are more operationally important for voice and video applications:

  • Audio history: Up to 100 turns and 600 seconds.
  • Video history: Up to 50 turns and 240 seconds.

When cumulative limits are exceeded, older history is discarded. An application that needs a long-running meeting or surveillance-style session should therefore maintain its own summaries, transcripts, or state outside the model and selectively send relevant context back into the active session.

The maximum output figure is unusually large for an interactive model, but real-time voice interfaces generally benefit from short, timely responses rather than long monologues. The 65,536-token value is a ceiling, not a recommendation for conversational response length.

Tool calling and real-time integration

The model supports custom function calling and remote MCP tools. Function calling allows the model to request an action from an application, such as looking up an order, controlling a permitted device, or retrieving information from a business system. The application remains responsible for executing the function and enforcing authentication, permissions, and safety checks.

Remote MCP support provides another way to connect the model to external tools and services. Because this model is designed for live interaction, tool latency becomes part of the user experience. A slow external service can make a voice conversation feel delayed even if the model itself responds quickly.

Streaming is supported, and the documented access methods include WebSocket, WebRTC, and AOQ. These options make the model suitable for systems that need incremental audio or text handling instead of waiting for the complete response before displaying or playing anything.

The supplied research does not verify fine-tuning, caching, batch API access, structured output, or a dedicated JSON mode for this model. Those capabilities should not be assumed merely because the model supports function calling.

Pricing and regional availability

Pricing is based on token usage and varies by region and modality. The supplied official pricing information lists the following rates per one million tokens:

DeploymentInputOutput
China (Beijing), text, images, and videoCNY 1.5Text: CNY 4.5
China (Beijing), audioCNY 6Audio: CNY 12
Singapore, text, images, and videoCNY 1.677Text: CNY 5.104
Singapore, audioCNY 6.781Audio: CNY 13.636

Speech output can incur both the corresponding text-output and audio-output charges. For example, an application that requests a spoken answer should account for the generated text as well as the audio representation, rather than budgeting only for one output category.

A one-million-token free quota is available in China (Beijing) for 90 days, subject to Alibaba Cloud's eligibility rules. The supplied research does not extend that free-quota offer to Singapore. The official model page lists both China (Beijing) and Singapore as supported regions, but regional access, account eligibility, quotas, and pricing should be checked before deployment.

For the international deployment, the supplied rate-limit documentation lists 60 requests per minute and 2,000,000 tokens per minute. These limits can affect concurrency planning for call-center, meeting, or multi-user applications.

Reasoning, coding, speed, and cost trade-offs

The supplied editorial evaluation rates the model's reasoning at 7 out of 10, coding at 5 out of 10, speed at 9 out of 10, and cost at 8 out of 10. These are comparative editorial estimates, not Alibaba Cloud benchmark results or provider-published scores.

The scores indicate the model's intended balance: very strong responsiveness and relatively favorable cost for a multimodal real-time model, with useful but not exceptional positioning for complex reasoning or software development. Its value is less about winning a general coding or long-form reasoning contest and more about responding quickly while processing speech and video.

For a voice assistant, fast turn-taking may matter more than maximum reasoning depth. For a code-generation workflow, however, a specialized coding model or a model optimized for extended software-engineering tasks may be a better choice. The supplied research does not identify a specific sibling model for that comparison, so no named alternative is asserted here.

Main strengths and limitations

Strengths

  • Designed specifically for real-time audio and video interaction.
  • Produces both text and spoken audio.
  • Supports streaming and multiple real-time transport options.
  • Can combine multimodal conversation with custom functions and remote MCP tools.
  • Supports multichannel audio through WebSocket.
  • Offers a large 196,608-token context and a 65,536-token maximum output.
  • Has documented regional pricing and an eligible China free-quota offer.

Limitations

  • Audio and video history have explicit turn and duration limits, after which older history is discarded.
  • Availability is limited to the documented supported regions and may vary by account or deployment.
  • Audio output can create two categories of output charges.
  • It is not an image, video, music, or embedding generation model.
  • The supplied research does not verify fine-tuning, caching, batch processing, structured output, or JSON mode.
  • Its real-time focus may make it less suitable than a specialized model for deep coding, batch analysis, or long offline workflows.

Best use cases

Qwen3.8-Omni-Flash-Realtime is a strong fit when the application must react to live audio or video and deliver a spoken response. Suitable examples include:

  • Voice assistants that need to answer while a conversation is in progress.
  • Speech-to-speech customer-service interfaces.
  • Interactive video agents that interpret a live or recent visual stream.
  • Meeting and collaboration tools that need audio-aware interaction.
  • Live media analysis with tool access for retrieving or updating external information.
  • Multimodal applications that use function calling or MCP to connect conversation with business systems.

For these uses, the combination of low-latency streaming, audio output, video understanding, and tool support is more important than a standalone text benchmark score.

When to choose this model

Choose Qwen3.8-Omni-Flash-Realtime when live turn-taking is central to the product and the system needs to combine speech, video, and actions. It is particularly compelling when users should be able to speak naturally, receive spoken answers, and invoke tools without leaving the conversation.

Consider another type of model when the workload is primarily offline text generation, large-scale batch processing, image or video creation, embedding generation, or conventional fine-tuning. A specialized coding or reasoning model may also be preferable when response speed is less important than sustained programming or complex analytical depth.

Before selecting it for production, verify the target region, transport method, rate limits, audio and video history behavior, and expected cost of dual text-and-audio output. The model's strongest distinction is its real-time multimodal interaction loop; applications that do not need that loop may not benefit from paying for those capabilities.

Bottom line

Qwen3.8-Omni-Flash-Realtime is a real-time speech-and-video model for Alibaba Cloud Model Studio, with text and audio output, streaming interfaces, tool calling, and remote MCP support. Its large context and fast-response positioning make it useful for live assistants and interactive multimodal systems. Its main constraints are regional availability, modality-specific pricing, bounded audio and video history, and a focus on real-time interaction rather than image generation, batch workloads, or specialized coding.


Answers to Frequently Asked Questions

How much does Qwen3.8-Omni-Flash-Realtime cost and where is it available?
The model is listed for China (Beijing) and Singapore, with international deployment available in Singapore. Pricing varies by region and modality, from CNY 1.5 per million input text, image, or video tokens in Beijing to CNY 6.781 per million input audio tokens in Singapore. Spoken responses can incur both text-output and audio-output charges. A China-based one-million-token free quota may be available for 90 days subject to eligibility rules.
What are the context and history limits of Qwen3.8-Omni-Flash-Realtime?
The model has a documented context length of 196,608 tokens and a maximum output length of 65,536 tokens. Real-time audio history is limited to 100 turns or 600 seconds, while video history is limited to 50 turns or 240 seconds. When these cumulative limits are exceeded, older history is discarded.
Which real-time interfaces and tools does Qwen3.8-Omni-Flash-Realtime support?
Alibaba Cloud Model Studio provides the model through WebSocket, WebRTC, and AOQ interfaces. It supports streaming, multichannel audio through WebSocket, custom function calling, and remote MCP tools. Applications remain responsible for executing functions and enforcing authentication, permissions, and safety controls.
What is Qwen3.8-Omni-Flash-Realtime used for?
Qwen3.8-Omni-Flash-Realtime is designed for real-time multimodal applications that process streaming audio, video frames, and text while a session is ongoing. Common use cases include speech-to-speech assistants, live video agents, meeting interfaces, customer-service tools, and applications that connect conversations with external business systems.
What input and output modalities does Qwen3.8-Omni-Flash-Realtime support?
The model supports text, streaming audio, and video represented as consecutive image frames as inputs. It can generate text and spoken audio as outputs. It does not generate images, videos, music, or embeddings.


Sources 5
Provider

About Qwen