Qwen-Audio-3.0-Realtime

qwen-audio-3.0-realtime-flash

by Qwen · Available; current API access confirmed, but newer models are recommended for some new projects

Alibaba Cloud's Qwen-Audio-3.0-Realtime-Flash is designed for duplex voice conversations with audio and text input/output. It supports streaming, interruption handling, function calling, optional web search, selectable voices and voice cloning, with a 40,960-token context window and 8,192-token maximum output. The model is available in China (Beijing) and Singapore through WebSocket, WebRTC and AOQ, with separate text and audio usage prices.

Text Speech Reasoning Coding
Qwen-Audio-3.0-Realtime-Flash is a currently accessible Alibaba Cloud Model Studio model designed for fast, streaming voice interaction. It handles audio and text input, can produce audio and text simultaneously, and supports a 40,960-token context window with up to 8,192 output tokens. Its main appeal is responsive duplex conversation rather than general-purpose reasoning, image processing or software development.
Outputs

What qwen-audio-3.0-realtime-flash can produce

Text Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
2/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Audio-3.0-Realtime
Model type Multimodal
Context window 41K tokens
Maximum output 8K tokens
Status Available; current API access confirmed, but newer models are recommended for some new projects
Knowledge cutoff notes

No authoritative provider-published knowledge cutoff was found for this exact real-time speech model. Optional web search can provide current external information during a session, but it does not establish or change the underlying model cutoff.

Model notes

The exact model accepts audio and text and returns audio and text. Output modalities can be configured as audio plus text by default or text only for debugging and logging. Supported system voices include longanqian, longanlingxin, longanlingxi, longanxiaoxin and longanlufeng; cloned voices are also supported. Input audio is documented as 16 kHz, 16-bit, mono PCM, while output audio is 24 kHz, 16-bit, mono PCM. Function calling and optional first-party web search are supported. The model supports WebSocket, WebRTC and AOQ transport options and is accessible in China (Beijing) and Singapore. The current international prices reflect Alibaba Cloud's September 22, 2026 price reduction; the model-specific page also displays older original prices. Editorial scores are comparative estimates, not provider benchmarks. Alibaba Cloud currently recommends newer speech-to-speech or omni models for some new projects, but the exact qwen-audio-3.0-realtime-flash model remains listed as accessible.

Cost

Model pricing

Input International/Singapore: text input $0.23 per 1M tokens; audio input $0.93 per 1M tokens. China (Beijing): text input $0.413 per 1M tokens; audio input $4.126 per 1M tokens.
Output International/Singapore: text-only output $0.70 per 1M tokens; combined text and audio output $1.87 per 1M tokens. China (Beijing): text-only output $4.126 per 1M tokens; combined text and audio output $13.752 per 1M tokens.
Model guide

Qwen-Audio-3.0-Realtime-Flash: Low-Latency Duplex Voice Conversations

Qwen-Audio-3.0-Realtime-Flash is Alibaba Cloud's real-time speech-to-speech model for low-latency, two-way voice conversations. It accepts audio and text, returns audio and text, supports streaming, interruption handling, function calling, optional web search, voice selection and voice cloning, and is available in selected Alibaba Cloud Model Studio regions.

What is Qwen-Audio-3.0-Realtime-Flash?

Qwen-Audio-3.0-Realtime-Flash is Alibaba Cloud's low-latency speech-to-speech model for applications that need a live voice exchange. “Duplex” means that the system is designed for two-way conversation: it can receive streaming speech while producing a spoken response, rather than waiting for a complete recording before generating an answer.

The model is part of the Qwen-Audio-3.0-Realtime family and is accessed through Alibaba Cloud Model Studio. The exact model is currently listed as available in the China (Beijing) and Singapore regions. Applications can connect through WebSocket, WebRTC or AOQ, depending on the integration approach.

This is not primarily a text chatbot with an optional voice interface. Its distinguishing purpose is real-time audio interaction, including voice assistants, customer-service agents and other experiences where response delay and interruption handling matter.

Audio, text and real-time conversation features

The verified input and output modalities are:

CapabilitySupport
Text inputYes
Audio inputYes
Text outputYes
Audio outputYes
Image input or outputNo
Video input or outputNo

Audio input is documented as 16 kHz, 16-bit, mono PCM. Generated audio is documented as 24 kHz, 16-bit, mono PCM. These format details are important when connecting microphones, telephony systems or media pipelines because an application may need to resample or convert audio before sending it.

The model can return audio and text together, while text-only output is available for debugging or logging. Supported system voices include longanqian, longanlingxin, longanlingxi, longanxiaoxin and longanlufeng. Voice cloning is also supported according to the supplied Model Studio documentation, although developers should separately verify the applicable permissions, consent requirements and regional policies before using cloned voices.

Real-time interaction features include streaming responses, interruption handling and conversation controls. These features are more relevant to a natural spoken exchange than to ordinary request-and-response text generation: a user can interrupt an answer, and the application can manage the continuing session rather than treating every utterance as an isolated recording.

Context window and output limits

Qwen-Audio-3.0-Realtime-Flash has a documented context length of 40,960 tokens and a maximum output of 8,192 tokens. A token is a unit used by the model to process text; audio is priced and handled separately according to the provider's audio-token accounting. The context limit covers the conversation and supplied content that the service retains for the request or session, so long-running voice applications should still manage history rather than assuming unlimited conversation memory.

The 8,192-token maximum is a generation ceiling, not a promise that every spoken response will be that long. For voice interfaces, shorter responses are generally more practical because they reduce the time before the user can act or speak again. The model's value comes from streaming and responsiveness, not from producing lengthy monologues.

Function calling, web search and reasoning

The model supports function calling. This allows an application to expose defined actions, such as checking an order, looking up an appointment or retrieving account information. The model can decide when a declared function is relevant and provide arguments for the application to execute. The external system, not the model itself, remains responsible for authorization, validation and carrying out the action.

Optional first-party web search is also supported. Web search can help a voice agent obtain more current external information, but it should not be confused with a change to the model's underlying knowledge cutoff. No authoritative knowledge cutoff was found for this exact real-time model.

Reasoning is available in the broad sense that the model can interpret a spoken request, maintain conversational context and select a tool when appropriate. However, the supplied research does not provide a formal reasoning benchmark or a specialized reasoning mode. The editorial reasoning score is 6 out of 10, which is a comparative assessment rather than an Alibaba Cloud-published benchmark.

Coding is not a primary strength. The editorial coding score is 2 out of 10, reflecting the model's focus on voice interaction rather than software development. It may help an agent perform a narrowly defined coding-related action through function calling, but a general-purpose coding model is more appropriate for substantial code generation or debugging.

Pricing by region and modality

Alibaba Cloud lists separate prices for Singapore or international service and China (Beijing). The billing unit is one million tokens, and audio and text are priced differently. The following figures are the currently supplied listed prices:

RegionInputOutput
International / SingaporeText input: $0.23 per 1M tokens
Audio input: $0.93 per 1M tokens
Text only: $0.70 per 1M tokens
Combined text and audio: $1.87 per 1M tokens
China (Beijing)Text input: $0.413 per 1M tokens
Audio input: $4.126 per 1M tokens
Text only: $4.126 per 1M tokens
Combined text and audio: $13.752 per 1M tokens

These prices are usage-based rather than a monthly consumer subscription. The amount an application spends depends on the duration and content of the conversation, the number of turns retained in context, whether audio or text output is requested, and the region used. Audio output is substantially more expensive than text-only output in both listed regional schedules, so an application that needs spoken responses should include that cost in its design.

The research notes that the international prices reflect a provider price reduction and that the model-specific page may also show older original prices. Developers should confirm the active price in the billing console before deployment, especially because Alibaba Cloud can update regional pricing and availability.

Main strengths and limitations

Where the model is strong

  • Low-latency voice interaction: Streaming audio responses and duplex conversation make it suitable for live exchanges.
  • Audio-native interaction: The model accepts speech and returns speech instead of requiring a separate speech-recognition and speech-synthesis pipeline for the core conversation.
  • Operational agent features: Function calling, interruption handling and optional web search support practical voice-agent workflows.
  • Voice customization: Multiple system voices and voice cloning provide options for branded or specialized experiences.
  • Multiple connection methods: WebSocket, WebRTC and AOQ give developers several ways to connect real-time clients and services.
  • Potentially efficient international pricing: The listed Singapore or international rates are considerably lower than the China (Beijing) rates for several input and output categories.

Limitations to plan for

  • Restricted modality scope: The model does not support image or video input or output, so it is not suitable for visual assistants or media-generation workflows.
  • Regional availability: The exact model is listed for China (Beijing) and Singapore, and access, account eligibility and quotas can change.
  • Audio cost: Audio input and especially combined audio-text output cost more than text-only usage in the listed pricing schedules.
  • Not a coding specialist: Its low editorial coding assessment reflects its intended role as a conversational speech model.
  • No verified structured-output support: The supplied research marks structured output as unsupported and JSON mode as unverified. Function calling should not automatically be treated as unrestricted JSON generation.
  • No fine-tuning, batch or caching support: The supplied model information lists fine-tuning, batch API and caching as unavailable.

Speed, cost and capability trade-offs

The editorial speed score is 9 out of 10 and the cost score is 8 out of 10. These are comparative editorial evaluations, not provider benchmarks. They reflect the model's real-time design and the relatively favorable international pricing, but they do not guarantee a particular latency, throughput or total bill.

For a voice assistant, a fast specialized model can be preferable to a larger general-purpose model that produces more elaborate reasoning but responds too slowly for natural conversation. Qwen-Audio-3.0-Realtime-Flash is therefore best understood as a task-focused compromise: it prioritizes responsiveness, streaming audio and conversation control over broad visual capability, advanced coding or maximum reasoning depth.

Cost also depends heavily on the output modality. Text-only responses can be useful for logging, testing or applications that perform speech synthesis elsewhere. When the model must generate audio itself, the combined text-and-audio output rate applies, and the application should account for the higher usage cost.

When to choose this model

Choose Qwen-Audio-3.0-Realtime-Flash when the central product requirement is a responsive spoken conversation. Suitable examples include:

  • Voice assistants that need to listen and respond continuously.
  • Real-time customer-service or support agents.
  • Interactive voice agents that call business functions such as appointment lookup or order status.
  • Hands-free interfaces where users cannot or do not want to type.
  • Streaming conversational prototypes that need selectable voices or voice cloning.
  • Applications where Singapore or China (Beijing) deployment is available and the regional pricing fits the budget.

It is less suitable when the main requirement is image understanding, video processing, image or video generation, large-scale batch inference, fine-tuning, structured data generation or advanced coding. For those tasks, a model with the relevant modality or specialization would be a better choice.

Position in Alibaba Cloud's current catalog

Qwen-Audio-3.0-Realtime-Flash remains listed as accessible, but Alibaba Cloud currently recommends newer speech-to-speech or omni models for some new projects. The supplied research specifically identifies Qwen-Audio-3.1-Realtime-Plus as a newer recommended speech-to-speech entry point and Qwen3.8-Omni-Flash-Realtime for newer real-time audio and video conversations. These recommendations do not make Qwen-Audio-3.0-Realtime-Flash unavailable; they indicate that developers starting a new project should compare the newer options against this model's documented access, pricing and feature requirements.

The practical reason to retain Qwen-Audio-3.0-Realtime-Flash on a shortlist is its clear focus on real-time audio conversation, established connection options and support for tools, voice choices and interruptions. The practical reason to investigate a newer sibling is the possibility of broader or more current capabilities, particularly when audio and video or newer speech-to-speech behavior is required.

Bottom line

Qwen-Audio-3.0-Realtime-Flash is a specialized real-time voice model rather than a general-purpose multimodal model. It combines audio and text input with audio and text output, streaming, interruption handling, function calling, optional web search and voice customization. Its strongest fit is a low-latency voice assistant or agent operating in a supported Alibaba Cloud region. Before choosing it, verify regional access, current prices, audio-token costs and whether the lack of visual, batch, fine-tuning and structured-output features creates a problem for the intended application.


Answers to Frequently Asked Questions

What are the main limitations of Qwen-Audio-3.0-Realtime-Flash?
The model does not support image or video input or output and is not designed for advanced coding, batch inference or fine-tuning. Structured output is listed as unsupported and JSON mode is unverified. Regional availability, audio costs and the lack of visual capabilities should be evaluated before deployment.
Where is Qwen-Audio-3.0-Realtime-Flash available and how much does it cost?
The model is currently listed in Alibaba Cloud Model Studio for the Singapore or international region and China (Beijing). Pricing is usage-based and differs by region and modality. Singapore or international rates range from $0.23 to $1.87 per 1 million tokens for the listed text and audio categories, while China (Beijing) rates range from $0.413 to $13.752 per 1 million tokens. Developers should confirm current prices and access in the billing console.
Does Qwen-Audio-3.0-Realtime-Flash support real-time interruptions and function calling?
Yes. The model supports streaming responses, interruption handling and function calling. Applications can expose actions such as checking an order or retrieving an appointment, while the external system remains responsible for authorization, validation and execution.
What is Qwen-Audio-3.0-Realtime-Flash best used for?
Qwen-Audio-3.0-Realtime-Flash is best suited to low-latency spoken interactions, including voice assistants, real-time customer-service agents, hands-free interfaces and voice agents that use business functions such as appointment or order lookups.
What audio formats does Qwen-Audio-3.0-Realtime-Flash support?
The model accepts 16 kHz, 16-bit, mono PCM audio input and generates 24 kHz, 16-bit, mono PCM audio output. Applications may need to resample or convert audio from microphones, telephony systems or other media pipelines.


Sources 8
Provider

About Qwen