Qwen-Audio-3.1-ASR

Qwen-Audio-3.1-ASR-Flash-Streaming

by Qwen · Current and available through Alibaba Cloud Model Studio in the International/Singapore and China (Beijing) regions.

A focused review of Alibaba Cloud Model Studio’s Qwen-Audio-3.1-ASR-Flash-Streaming model, covering its WebSocket streaming workflow, multilingual and Chinese-dialect recognition, timestamps, hot words, context biasing, token limits, regional pricing, and unsupported features such as diarization, emotion recognition, function calling, batch inference, and fine-tuning.

Text Reasoning Coding
Qwen-Audio-3.1-ASR-Flash-Streaming is designed for applications that need continuous speech-to-text results while audio is still being received. It can support live meeting transcription, captions, voice interfaces, and multilingual customer-service workflows through Alibaba Cloud Model Studio’s real-time WebSocket API. Its main advantage is focused, low-latency transcription rather than broad reasoning or tool use.
Outputs

What Qwen-Audio-3.1-ASR-Flash-Streaming can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Audio-3.1-ASR
Model type Speech Recognition
Context window 8K tokens
Maximum output 1K tokens
Status Current and available through Alibaba Cloud Model Studio in the International/Singapore and China (Beijing) regions.
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this dedicated streaming speech-recognition model.

Model notes

The exact model is a dedicated real-time ASR system rather than a general-purpose language model. It accepts audio and returns text through WebSocket. Alibaba Cloud documents hot words, prompt context, sentence-level and word-level timestamps, punctuation prediction, text normalization, Chinese-English switching, multilingual recognition, Chinese-dialect recognition, and noise robustness. The streaming model does not support speaker diarization or emotion recognition. Alibaba Cloud lists a 600-RPM limit for the International/Singapore and China (Beijing) deployments. Pricing is token-based for version 3.1; it should not be confused with the older Qwen-Audio-3.0 streaming model, which is priced by input-audio duration.

Cost

Model pricing

Input International/Singapore: USD 0.93 per 1 million input tokens; China (Beijing): USD 0.848 per 1 million input tokens.
Output International/Singapore: USD 0.70 per 1 million output tokens; China (Beijing): USD 0.636 per 1 million output tokens.
Model guide

Qwen-Audio-3.1-ASR-Flash-Streaming for Real-Time Multilingual Transcription

Qwen-Audio-3.1-ASR-Flash-Streaming is Alibaba Cloud Model Studio’s WebSocket-based speech-recognition model for turning live audio into text. It supports multilingual transcription, several Chinese dialects, hot words, contextual biasing, timestamps, punctuation, text normalization, and noise-robust recognition, but it does not provide speech synthesis, speaker diarization, emotion recognition, function calling, or general-purpose language-model features.

What Qwen-Audio-3.1-ASR-Flash-Streaming is

Qwen-Audio-3.1-ASR-Flash-Streaming is a dedicated automatic speech-recognition model from Alibaba Cloud Model Studio. Automatic speech recognition, or ASR, converts spoken audio into written text. Unlike a general-purpose conversational model, this model is focused on recognizing speech as it arrives and returning transcription events during an active streaming session.

The model is accessed through a WebSocket real-time speech-recognition API. WebSocket connections allow the application and service to keep an ongoing two-way connection open, so audio can be sent in chunks instead of waiting for a complete recording to finish. The service can then return partial and completed recognition results as the conversation progresses.

Alibaba Cloud positions the model for live transcription, conferences, broadcasts, subtitles, voice commands, and intelligent voice interaction. It belongs to the Qwen-Audio-3.1 ASR family, but the streaming version should be treated as a specialized speech-recognition endpoint rather than as a general chat or reasoning model.

Core capabilities and supported audio workflows

The model accepts audio and produces text. Supported streaming input configurations include PCM and Opus audio, with documented sample rates of 16 kHz and 8 kHz; 16 kHz is the default listed by the real-time API documentation. Audio is sent as Base64-encoded chunks through WebSocket events.

Its documented capabilities include:

  • Real-time transcription while audio is being transmitted.
  • Multilingual speech recognition across Chinese, English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak, and Norwegian.
  • Recognition of several Chinese regional varieties, including Shanghai, Nanchang, Ningbo, Hakka, Hangzhou, Wenzhou, Hunan, Fujian, Cantonese, and Suzhou dialects.
  • Chinese-English switching within the same recognition workflow.
  • Hot-word support for names, product terms, specialist vocabulary, and other words that may otherwise be difficult to recognize.
  • Prompt or contextual biasing to provide additional recognition context.
  • Sentence-level and word-level timestamps for caption alignment and keyword highlighting.
  • Punctuation prediction and text normalization.
  • Noise robustness for more complex acoustic environments.

Hot words and contextual biasing are particularly useful when an application has a known vocabulary. For example, a meeting application could provide company names or technical terms that are likely to appear in the discussion. These features can help guide recognition, but they should not be treated as a guarantee that every specialized term will be transcribed correctly.

Context, output limits, and API details

The model documentation lists an 8,192-token context window, a maximum input length of 8,192 tokens, and a maximum output length of 1,024 tokens. These are token limits in the model documentation and should not be confused with a fixed maximum recording duration. Alibaba Cloud’s ASR overview lists the maximum audio duration for streaming use as unlimited.

“Unlimited” streaming duration does not mean that every connection or application is free from operational constraints. Long-running sessions still need appropriate client-side handling for connection failures, partial results, completed utterances, and service rate limits. The available research specifically documents a 600-requests-per-minute limit for the International deployment in Singapore. Applications should verify the current regional documentation before designing capacity assumptions.

The streaming interface is intended for continuous recognition rather than offline file-processing workflows. A system that receives completed recordings, needs asynchronous processing, or requires additional speech-analysis features may be better served by an appropriate file-transcription or specialized speech model instead.

Pricing and regional availability

Qwen-Audio-3.1-ASR-Flash-Streaming is listed for the International deployment scope in Singapore and for the China (Beijing) region. Pricing is token-based for this version:

DeploymentInput priceOutput price
International, SingaporeUSD 0.93 per 1 million input tokensUSD 0.70 per 1 million output tokens
China, BeijingUSD 0.848 per 1 million input tokensUSD 0.636 per 1 million output tokens

The figures above are the documented regional prices supplied for this model. They should not be confused with the older Qwen-Audio-3.0 streaming model, which is described as using input-audio-duration pricing. Actual charges, eligibility, and regional availability can change, so production deployments should confirm the current Alibaba Cloud Model Studio pricing page and the selected region.

Important limitations and unsupported features

This model is deliberately narrow in scope. Alibaba Cloud’s current documentation marks the following features as unsupported for the exact streaming model:

  • Function calling and tool invocation.
  • Structured outputs.
  • Web search.
  • Context caching.
  • Batch inference.
  • Fine-tuning.
  • Speaker diarization.
  • Emotion recognition.

Speaker diarization identifies which person is speaking. Because it is not supported here, a transcript from a meeting should not be expected to contain reliable speaker labels such as “Speaker 1” and “Speaker 2.” Emotion recognition is also unavailable, so the model should not be used as a built-in detector of mood or sentiment from voice.

The output is text only. The model does not synthesize speech, generate music, return images or video, or provide other non-text media. It also is not intended for general conversation, coding, complex reasoning, or autonomous agent tasks. The absence of function calling and web search means that an application cannot use this endpoint as a complete voice agent without adding separate components for dialogue, tools, retrieval, and response generation.

Reasoning, coding, and tool-use profile

Qwen-Audio-3.1-ASR-Flash-Streaming should not be evaluated like a general large language model. Its job is to recognize and format spoken language, not to solve multi-step problems or write software. The supplied specifications do not document a general reasoning mode or coding capability for this model.

Likewise, tool use is not part of the endpoint’s documented feature set. Function calling, structured output, and web search are listed as unsupported. If a voice application needs to turn recognized speech into actions, a typical architecture would use this model for transcription and a separate application layer or language model for intent handling and tool execution. That separation is an architectural recommendation, not a built-in capability of this model.

Speed, cost, and capability trade-offs

The model’s main trade-off is specialization. It concentrates its capabilities on streaming speech recognition, which makes it a more appropriate choice for live captions or voice input than a general multimodal model that is designed to understand images, reason over documents, or call tools. The supplied evaluation data gives it a high editorial speed assessment and a relatively favorable editorial cost assessment, but those are database evaluations rather than Alibaba Cloud benchmark claims.

Its focused design also means that it lacks the broader feature set of a general-purpose model. A general language model may be more suitable when the application needs summarization, question answering, coding, planning, or function execution after speech is received. A speech model with diarization may be preferable for interview or meeting transcripts where identifying each participant is essential. A file-transcription service may be more appropriate when the input is a completed recording rather than a live stream.

Best use cases

  • Live meeting transcription: stream microphone audio and display partial and completed text during a meeting.
  • Real-time subtitles: use sentence and word timestamps to align recognition results with a broadcast, presentation, or live event.
  • Multilingual voice interfaces: accept speech in supported languages before passing the transcript to a separate dialogue or application system.
  • Customer-service assistance: transcribe ongoing calls or voice interactions, especially when domain terms can be supplied as hot words or contextual prompts.
  • Voice commands: convert spoken commands into text for an application layer that performs its own intent detection and authorization.
  • Keyword highlighting: use word-level timestamps to identify and highlight terms in a live transcript.
  • Chinese-dialect applications: support workflows involving the documented Chinese regional varieties when the deployment and language requirements match.

When to choose this model

Choose Qwen-Audio-3.1-ASR-Flash-Streaming when the central requirement is continuous, low-latency transcription over WebSocket. It is a strong fit when an application needs multilingual recognition, Chinese-dialect coverage, timestamps, punctuation, hot words, and contextual biasing without requiring the model itself to reason, browse, or execute tools.

Choose another option when the workflow depends on speaker labels, emotion analysis, offline file processing, or speech synthesis. A diarization-capable transcription model is more suitable for multi-speaker records. A file-oriented ASR service is more suitable for batch uploads and asynchronous processing. A general language model or voice-agent architecture is more suitable when transcription is only the first step and the system must answer questions, summarize discussions, write code, call business tools, or take actions.

Bottom line

Qwen-Audio-3.1-ASR-Flash-Streaming is a focused real-time transcription model rather than an all-purpose AI assistant. Its documented strengths are streaming delivery, broad language coverage, Chinese-dialect recognition, timestamps, contextual vocabulary support, punctuation, and noise robustness. Its documented boundaries are equally important: text output only, no diarization or emotion recognition, no function calling or web search, and no batch or fine-tuning support. For live speech-to-text applications that fit those boundaries, it offers a clear and technically specific role in Alibaba Cloud Model Studio’s current ASR catalog.


Answers to Frequently Asked Questions

What is Qwen-Audio-3.1-ASR-Flash-Streaming used for?
Qwen-Audio-3.1-ASR-Flash-Streaming is designed for real-time speech-to-text transcription over a WebSocket connection. Common uses include live meeting transcription, subtitles, broadcasts, multilingual voice interfaces, customer-service assistance, voice commands, and keyword highlighting.
Which languages and audio formats does Qwen-Audio-3.1-ASR-Flash-Streaming support?
The model supports many languages, including Chinese, English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak, and Norwegian. It also supports several Chinese regional varieties, Chinese-English switching, PCM and Opus audio, and documented 8 kHz and 16 kHz sample rates.
Can Qwen-Audio-3.1-ASR-Flash-Streaming call tools or perform general AI tasks?
No. Function calling, structured outputs, web search, coding, complex reasoning, and general conversational capabilities are not supported as built-in features. Applications that need intent detection, dialogue, retrieval, or tool execution should use the transcription output with a separate application layer or language model.
How much does Qwen-Audio-3.1-ASR-Flash-Streaming cost and where is it available?
The model is listed for the International deployment in Singapore and the China (Beijing) region. Singapore pricing is USD 0.93 per 1 million input tokens and USD 0.70 per 1 million output tokens. Beijing pricing is USD 0.848 per 1 million input tokens and USD 0.636 per 1 million output tokens. Pricing, eligibility, and regional availability should be confirmed in the current Alibaba Cloud Model Studio documentation.
Does Qwen-Audio-3.1-ASR-Flash-Streaming support speaker diarization or emotion recognition?
No. The exact streaming model does not support speaker diarization or emotion recognition. It should not be expected to reliably label different speakers or detect mood and sentiment from voice.


Sources 6
Provider

About Qwen