Qwen-Audio

qwen-audio-3.1-realtime-plus

by Qwen · Current and available

A real-time duplex speech model from Alibaba Cloud Model Studio that supports streaming text and audio interaction, transcripts, spoken responses, function calling, optional web search, configurable voices, and voice cloning.

Text Speech Reasoning Coding
Qwen-Audio 3.1 Realtime Plus is built for applications that need a live spoken conversation rather than a sequence of separate speech-recognition and text-to-speech services. It can receive streaming microphone audio, generate spoken responses and transcripts, detect conversational turns, and connect to external tools through function calling.
Outputs

What qwen-audio-3.1-realtime-plus can produce

Text Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

4/10 Reasoning
3/10 Coding
9/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Audio
Model type Multimodal
Context window 262K tokens
Maximum output 16K tokens
Release date 2026-09-20
Status Current and available
Knowledge cutoff notes

Alibaba Cloud's model documentation does not publish a specific knowledge-cutoff date for this real-time speech model.

Model notes

Real-time duplex speech model with text and audio input/output. Supports WebSocket, WebRTC, and AOQ access. The default voice is longanqian_v3.1; additional system voices and cloned voices are supported. Audio input is 16 kHz, 16-bit, mono PCM, while audio output is 24 kHz, 16-bit, mono PCM. Supports acoustic VAD, semantic turn detection, and push-to-talk modes. Function calling and web search cannot be enabled together. The model supports up to 50 audio turns or 300 seconds of retained audio history. Listed rate limits are 60 RPM and 100,000 TPM.

Cost

Model pricing

Input Singapore: $0.80 per 1M text-input tokens; $6.40 per 1M audio-input tokens. China (Beijing): $0.688 per 1M text-input tokens; $5.501 per 1M audio-input tokens.
Output Singapore: $6.40 per 1M text-output tokens; $24 per 1M audio-output tokens. China (Beijing): $5.501 per 1M text-output tokens; $20.628 per 1M audio-output tokens.
Model guide

Qwen-Audio 3.1 Realtime Plus for Full-Duplex Voice Conversations

Qwen-Audio 3.1 Realtime Plus is Alibaba Cloud Model Studio's real-time duplex speech model for low-latency voice applications. It accepts text and audio, produces text and audio, supports streaming interaction, function calling, optional web search, voice cloning, and a 262,144-token context window.

What is Qwen-Audio 3.1 Realtime Plus?

Qwen-Audio 3.1 Realtime Plus is a real-time speech model provided through Alibaba Cloud Model Studio. Its main purpose is to support natural, low-latency voice interaction. Instead of requiring an application to send audio to one service for transcription, pass the resulting text to a language model, and then send the answer to a text-to-speech system, this model can handle the spoken interaction as a streaming session.

The model accepts both text and audio input and can return both text and audio output. That makes it suitable for voice assistants, customer-service systems, call-center interfaces, AI companions, and hands-free applications. The model is part of the Qwen-Audio family and is specifically positioned for real-time duplex communication, where audio can be sent and received continuously rather than only in isolated request-and-response turns.

How the real-time interaction works

Qwen-Audio 3.1 Realtime Plus supports full-duplex streaming. In practical terms, an application can continue sending microphone audio while the model generates response audio and text incrementally. This can reduce the delay between a user's speech and the beginning of the assistant's reply, provided that the client handles audio buffering and playback correctly.

The model supports connections through WebSocket, WebRTC, and AOQ. Alibaba Cloud documents server-side acoustic voice activity detection, semantic turn detection, and push-to-talk modes. Acoustic voice activity detection helps identify when speech is present, while semantic turn detection uses the meaning and conversational context of the input to help determine whether the speaker has finished. Push-to-talk can be preferable when an application needs explicit control over when audio is submitted.

Applications must still manage interruption behavior, audio buffering, playback, and turn handling. A real-time model does not automatically remove those engineering responsibilities. For the documented format, input audio uses 16 kHz, 16-bit, mono PCM, while output audio uses 24 kHz, 16-bit, mono PCM.

Inputs, outputs, and supported capabilities

CapabilityDocumented support
Text inputYes
Audio inputYes, 16 kHz, 16-bit, mono PCM
Text outputYes, including incremental transcripts
Audio outputYes, 24 kHz, 16-bit, mono PCM
StreamingYes
Function callingYes
Web searchOptional, using the enable_search setting
Image or video inputNo documented support for this model
Image or video outputNo

Function calling lets the model request an action from the surrounding application, such as looking up an account, checking an order, or controlling a connected workflow. The application, rather than the model itself, executes the function and returns the result to the conversation. Qwen-Audio 3.1 Realtime Plus also supports an optional first-party web-search capability. Function calling and web search cannot be enabled at the same time in a session, so an implementation must choose which tool path it needs.

The model supports multiple system voices, including the documented default voice longanqian_v3.1. Alibaba Cloud's voice-cloning feature can also provide cloned voices. Voice selection and cloning should be treated as product and policy decisions as well as technical features, particularly when an application uses a recognizable person's voice or processes voice data from users.

Context window and conversation history

The documented context window is 262,144 tokens. The maximum input length is 245,760 tokens, and the maximum output length is 16,384 tokens. These limits are substantially larger than what is normally needed for a short spoken exchange, but they can help applications retain extensive instructions, transcripts, tool results, or long-running conversational context.

Audio history has a separate practical limit: the model supports up to 50 audio turns or 300 seconds of retained audio history. This distinction matters because a large token context window does not mean that unlimited raw audio can remain active in a session. Applications with long conversations may need to summarize, trim, or otherwise manage earlier context.

Pricing and regional access

Alibaba Cloud Model Studio prices this model by token volume, with separate rates for text and audio and different prices by deployment region. In Singapore, the listed prices are:

UsagePrice per 1 million tokens
Text input$0.80
Audio input$6.40
Text output$6.40
Audio output$24.00

For the China (Beijing) deployment, the listed rates are $0.688 per million text-input tokens, $5.501 per million audio-input tokens, $5.501 per million text-output tokens, and $20.628 per million audio-output tokens. These are usage prices rather than a consumer subscription fee. Actual costs depend on how much audio and text an application sends and receives, how much history it retains, and which region it uses.

The documented rate limit is 60 requests per minute and 100,000 tokens per minute. Availability, quotas, account requirements, and regional access should be confirmed in the relevant Alibaba Cloud deployment before production use.

Speed, cost, and capability trade-offs

The model's main advantage is its real-time speech workflow. It is designed to begin producing a response while an interaction is still being streamed, which is more appropriate for live conversation than a conventional batch-oriented language model. The model also combines spoken input, spoken output, turn detection, voices, and tool features in one real-time interface.

Audio tokens are more expensive than text tokens in the listed Singapore pricing. Audio output is particularly costly at $24 per million tokens, compared with $6.40 for text output. This creates a clear cost trade-off: applications that need spoken responses should budget for audio-output usage, while applications that only need transcription or text responses may be able to reduce costs by limiting audio generation. Streaming can improve perceived responsiveness, but it does not by itself make every interaction cheaper.

The supplied research does not provide benchmark results or a provider-published reasoning score for this model. Its documented strengths are centered on speech interaction and latency rather than general-purpose reasoning or coding performance. Function calling can extend an application with external actions, but that should not be interpreted as evidence that the model is optimized for complex software development or general-purpose analytical workloads.

Best use cases

  • Voice assistants: The model can listen to spoken requests, respond with synthesized speech, and maintain a live conversational flow.
  • Customer service: Full-duplex interaction, turn detection, transcripts, and function calling can support account, order, or support workflows when connected to application tools.
  • Call-center and voice-interface systems: Streaming audio and server-side turn handling are useful where response delay affects the user experience.
  • Hands-free applications: The model can power spoken interfaces for situations in which users cannot or do not want to interact with a screen.
  • AI companions and interactive characters: Multiple voices and cloned voices can support character-oriented experiences, subject to appropriate consent and usage controls.
  • Voice applications with current information: Optional web search may be useful when a conversation needs access to retrieved information rather than only the model's built-in knowledge.

When to choose Qwen-Audio 3.1 Realtime Plus

Choose this model when the central requirement is a live spoken conversation with low perceived latency. It is especially relevant when an application needs audio input and audio output in the same session, server-side turn detection, configurable voices, or real-time tool access. Its support for WebSocket, WebRTC, and AOQ gives teams several integration paths depending on their client and server architecture.

It may be a good fit when a unified speech-to-speech workflow is more useful than assembling separate speech-recognition, text-generation, and text-to-speech components. It can also be attractive when a voice application needs both transcripts and spoken output, since the model can return text as well as audio.

When another option may be more appropriate

A different model or architecture may be preferable for image or video understanding, image generation, video generation, embeddings, batch inference, or fine-tuning. Qwen-Audio 3.1 Realtime Plus is specialized for live audio interaction and is not documented as a general multimodal model for those workloads.

A text-only model may be more economical when users do not need spoken responses, because the listed audio-input and audio-output rates are higher than the corresponding text rates. A separate speech-recognition and text-to-speech pipeline may also be easier to customize when an application requires highly specific audio processing, independent vendor selection, or different voices and latency controls at each stage.

Applications that need both web search and custom function calling in the same session should also examine the design carefully, because the two features cannot be enabled together. If a workflow depends on several external systems, the application may need to implement its own retrieval or orchestration layer instead of relying on the model's optional search feature.

Limitations and implementation notes

The most important implementation limitation is that real-time voice quality depends on the surrounding client. Developers need to send audio in the expected format, process streamed response deltas, begin playback promptly, and handle interruptions or changes of turn. Input and output use different sample rates, so an audio device or client may need conversion.

The model's 50-turn or 300-second retained-audio limit also requires conversation-history management for long sessions. Large textual context capacity does not remove the need to control accumulated audio history. Finally, voice cloning and voice-based applications require careful handling of consent, privacy, and user expectations.

Overall, Qwen-Audio 3.1 Realtime Plus is best understood as a specialized real-time speech engine within Alibaba Cloud Model Studio. Its value comes from combining streaming duplex audio, text transcripts, configurable voices, turn detection, and optional tools. It is not the default choice for every AI workload, but it is a focused option for applications where the primary interface is a live spoken conversation.


Answers to Frequently Asked Questions

What are the main limitations of Qwen-Audio 3.1 Realtime Plus?
The model is specialized for live audio interaction rather than image, video, embeddings, batch inference, or fine-tuning workloads. It supports up to 50 audio turns or 300 seconds of retained audio history, requires careful client-side handling of buffering and interruptions, and cannot use web search and custom function calling simultaneously.
How much does Qwen-Audio 3.1 Realtime Plus cost?
In the Singapore region, the listed prices per 1 million tokens are $0.80 for text input, $6.40 for audio input, $6.40 for text output, and $24.00 for audio output. China (Beijing) has separate listed rates. Actual costs depend on audio and text volume, retained history, and deployment region.
What audio formats does Qwen-Audio 3.1 Realtime Plus support?
The documented input format is 16 kHz, 16-bit, mono PCM audio. The documented output format is 24 kHz, 16-bit, mono PCM audio. Applications may need to convert audio between these sample rates and manage buffering and playback.
Does Qwen-Audio 3.1 Realtime Plus support function calling and web search?
Yes, the model supports function calling and optional web search through the enable_search setting. However, function calling and web search cannot be enabled at the same time in one session, so an application must choose the appropriate tool path.
What is Qwen-Audio 3.1 Realtime Plus?
Qwen-Audio 3.1 Realtime Plus is a real-time speech model in Alibaba Cloud Model Studio that supports streaming audio and text input, along with audio and text output. It is designed for low-latency, full-duplex voice conversations in applications such as voice assistants, customer-service systems, call centers, and hands-free interfaces.


Sources 5
Provider

About Qwen