Qwen-Audio 3.1 ASR

Qwen-Audio-3.1-ASR-Flash-Filetrans

by Qwen · Current and publicly available

Alibaba Cloud's specialized asynchronous ASR model for transcribing long audio and video files. It supports multilingual speech, regional Chinese dialects, speaker separation, timestamps, hot words, context enhancement, punctuation, and text normalization, with documented limits of 2 GB and 12 hours.

Text Reasoning Coding
Qwen-Audio-3.1-ASR-Flash-Filetrans is built for offline transcription of long recordings rather than live conversations. It accepts a publicly accessible audio or video file URL, processes the recording asynchronously, and returns text transcription with options such as timestamps, speaker separation, hot-word enhancement, punctuation, and text normalization. It is suited to meetings, interviews, calls, subtitles, media archives, and other workflows where low-latency streaming is not required.
Outputs

What Qwen-Audio-3.1-ASR-Flash-Filetrans can produce

Text
Inputs

What it can understand

Audio
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Audio 3.1 ASR
Model type Other
Context window 8K tokens
Maximum output 1K tokens
Status Current and publicly available
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was published in the reviewed Alibaba Cloud documentation. This is a speech-recognition model, so a conventional text-model knowledge cutoff may not be applicable.

Model notes

The canonical API identifier is qwen-audio-3.1-asr-flash-filetrans. It is an asynchronous HTTP transcription model that accepts one publicly accessible HTTP or HTTPS audio or video file URL per request. Supported files can reach 2 GB and 12 hours, although recordings using speaker diarization should generally remain within two hours. The model supports timestamps, speaker separation, hot words, context enhancement, punctuation prediction, text normalization, multilingual recognition, and multiple regional Chinese dialects. It does not support emotion recognition according to Alibaba Cloud documentation. Pricing is token-based and region-dependent. The editorial scores emphasize transcription specialization rather than general reasoning or coding ability.

Cost

Model pricing

Input USD 0.15 per 1 million input tokens in Singapore; USD 0.113 per 1 million input tokens in China (Beijing)
Output USD 0.47 per 1 million output tokens in Singapore; USD 0.382 per 1 million output tokens in China (Beijing)
Model guide

Qwen-Audio-3.1-ASR-Flash-Filetrans for Long-Audio Transcription

Qwen-Audio-3.1-ASR-Flash-Filetrans is Alibaba Cloud Model Studio's asynchronous speech-recognition model for transcribing long audio and video files. It focuses on multilingual transcription, regional Chinese dialects, speaker separation, timestamps, punctuation, text normalization, hot words, and contextual recognition for recordings up to 2 GB or 12 hours.

What Qwen-Audio-3.1-ASR-Flash-Filetrans does

Qwen-Audio-3.1-ASR-Flash-Filetrans is Alibaba Cloud Model Studio's offline, end-to-end automatic speech recognition model. Automatic speech recognition, or ASR, converts spoken audio into written text. This particular model is aimed at long recordings rather than short conversational turns or real-time voice interaction.

The model's canonical API identifier is qwen-audio-3.1-asr-flash-filetrans. A request submits one publicly accessible HTTP or HTTPS URL for an audio or video file. Alibaba Cloud then processes the file asynchronously: the initial request returns a task identifier, and the application retrieves the completed transcription or receives it through the supported result workflow.

The model is available through Alibaba Cloud Model Studio in the Singapore international region and the China (Beijing) region. Its role in the Qwen-Audio 3.1 family is specialized: it is the file-based, non-real-time transcription option, while Alibaba Cloud documents a separate Qwen-Audio-3.1-ASR-Flash-Streaming model for real-time recognition.

Recognition capabilities

Its main advantage is the collection of controls designed for difficult or information-dense recordings. Documented capabilities include multilingual speech recognition, recognition of multiple regional Chinese dialects, punctuation prediction, text normalization, hot-word enhancement, context enhancement, and speaker separation.

  • Multilingual and dialect recognition: useful for recordings containing supported languages or regional Chinese speech varieties.
  • Speaker separation: distinguishes speakers in a multi-person recording, making meeting minutes, interviews, and call analysis easier to organize.
  • Timestamps: associates transcript content with points in the recording, which is useful for subtitles, searching, editing, and reviewing evidence.
  • Hot words: improves recognition of important names, brands, technical terms, or organization-specific vocabulary supplied as contextual terms.
  • Context enhancement: gives the recognizer additional information that can help resolve words that are difficult to identify from audio alone.
  • Punctuation and normalization: converts raw recognized speech into more readable text and applies text-formatting or normalization behavior supported by the service.
  • Noise robustness: Alibaba Cloud describes the model as suitable for complex acoustic environments, although the supplied documentation does not provide a benchmark score or guaranteed noise threshold.

These features are provider-documented capabilities, not independent benchmark results. Recognition quality will still depend on factors such as recording quality, overlapping speech, accents, vocabulary, and background noise.

File types, size limits, and processing behavior

The non-real-time API accepts one publicly accessible HTTP or HTTPS audio or video file URL per request. Supported formats listed by Alibaba Cloud include AAC, AMR, AVI, FLAC, FLV, M4A, MKV, MOV, MP3, MP4, MPEG, OGG, OPUS, WAV, WEBM, WMA, and WMV.

The maximum documented file size is 2 GB, and the maximum duration is 12 hours. Those limits make the model appropriate for long meetings, lectures, interviews, broadcasts, and archived media that would be inconvenient to divide into short clips. Alibaba Cloud recommends keeping recordings to roughly two hours when speaker diarization is enabled because longer diarized jobs have a greater risk of recognition failures or timeouts.

The URL requirement is an important practical constraint. The model does not provide a native local-file upload workflow in the documented request format. A file must first be placed at an address that the service can access over HTTP or HTTPS. Access permissions, URL expiration, and the availability of the media during processing therefore become part of the integration design.

Input, output, and API profile

The model accepts recorded audio and video files as input, but its result is text transcription. It does not generate audio, music, images, or video. A video file is used as a media container for its audio track; this is not documented as visual-video understanding.

SpecificationDocumented value
Model IDqwen-audio-3.1-asr-flash-filetrans
Primary taskAsynchronous, offline speech-to-text transcription
InputOne publicly accessible audio or video file URL
OutputText transcription, with supported timestamps and speaker information
Maximum file size2 GB
Maximum duration12 hours
Context length8,192
Maximum output tokens1,024
ProcessingAsynchronous; not real-time streaming
Documented regionsSingapore international and China (Beijing)

The listed 8,192 context length and 1,024 maximum output-token values are catalog specifications supplied for this model. They should not be interpreted as a promise that a 12-hour recording will be returned as one short response. Long-file transcription is handled by the service's asynchronous workflow and may involve result structures appropriate to the recording and enabled options.

The model is not documented as supporting function calling, structured outputs, web search, prefix completion, context caching, batch inference, or fine-tuning. It is therefore best treated as a focused transcription endpoint rather than a general-purpose assistant or agent.

Pricing and throughput

Alibaba Cloud lists token-based pricing for this model rather than a simple per-minute recording fee. In the Singapore international region, the listed price is USD 0.15 per 1 million input tokens and USD 0.47 per 1 million output tokens. In China (Beijing), the listed prices are USD 0.113 per 1 million input tokens and USD 0.382 per 1 million output tokens.

Because billing is token-based, the final cost depends on the service's token accounting for the submitted audio and generated transcription rather than only on the file's duration. The two regional price schedules should not be mixed: the applicable rate depends on where the request is processed.

Alibaba Cloud lists a limit of 600 requests per minute in both documented regions. That rate limit concerns request throughput, not a guarantee that every recording will finish within a particular time. Long files, speaker separation, file accessibility, and service conditions can affect completion time.

Main strengths and limitations

Where it is strong

  • It is designed specifically for long, asynchronous recordings rather than requiring applications to manage many short real-time chunks.
  • Speaker separation and timestamps support practical downstream tasks such as meeting minutes, subtitle preparation, interview review, and searchable archives.
  • Multilingual recognition and regional Chinese dialect support broaden its use beyond standard single-language recordings.
  • Hot words and context enhancement can help organizations handle names, specialist vocabulary, and recurring terminology.
  • The 2 GB and 12-hour limits accommodate many full-length recordings.

Important limitations

  • It is not a real-time transcription model. Applications needing live captions or immediate voice interaction should use a streaming ASR option instead.
  • Media must be available through a publicly accessible HTTP or HTTPS URL, which may require a separate storage and access-control step.
  • Speaker diarization is more demanding, and Alibaba Cloud recommends shorter recordings when that option is enabled.
  • There is no documented support for text-to-speech, audio generation, general conversation, image or video generation, tool use, or web search.
  • The supplied research does not provide an independent accuracy benchmark, a guaranteed word-error rate, or a documented emotion-recognition feature. Emotion recognition should not be assumed.
  • Token-based regional pricing may be less straightforward to forecast than a clearly stated per-minute transcription price.

When to choose this model

Choose Qwen-Audio-3.1-ASR-Flash-Filetrans when the central requirement is reliable batch processing of recorded speech and the application can tolerate asynchronous completion. Good examples include converting a two-hour meeting into a speaker-labeled transcript, indexing a collection of interviews, generating subtitle drafts from media files, transcribing multilingual customer calls, or making long audio archives searchable.

It is especially suitable when timestamps, dialect recognition, hot words, or speaker separation matter more than conversational reasoning. The model can provide the transcript, but a separate application or model may still be needed to summarize meetings, extract action items, classify calls, or answer questions about the resulting text.

Choose a streaming speech-recognition model instead when users need live captions, immediate partial transcripts, or interactive voice input. The related Qwen-Audio-3.1-ASR-Flash-Streaming model is the relevant sibling mentioned in Alibaba Cloud's documentation for real-time recognition. Choose a general-purpose language model after transcription when the primary task is analysis, writing, coding, or question answering rather than speech recognition itself.

Overall assessment

Qwen-Audio-3.1-ASR-Flash-Filetrans is a specialized long-file transcription service, not a broad multimodal assistant. Its value comes from combining long-duration asynchronous processing with practical ASR controls such as speaker separation, timestamps, hot words, context enhancement, punctuation, and text normalization. Its limitations are equally clear: it is not designed for live audio, it requires an accessible media URL, and its documented API does not include general tools or generation features.

For organizations processing recorded meetings, interviews, calls, subtitles, or media archives, the model offers a focused balance of scale and transcription features. The main decisions are whether asynchronous processing is acceptable, whether the file can be exposed through an accessible URL, and whether regional token pricing fits the workload.


Answers to Frequently Asked Questions

How does pricing work for Qwen-Audio-3.1-ASR-Flash-Filetrans?
Pricing is token-based rather than a simple per-minute fee. In the Singapore international region, Alibaba Cloud lists USD 0.15 per 1 million input tokens and USD 0.47 per 1 million output tokens. In China (Beijing), the listed prices are USD 0.113 per 1 million input tokens and USD 0.382 per 1 million output tokens.
Is Qwen-Audio-3.1-ASR-Flash-Filetrans a real-time transcription model?
No. It is an offline, asynchronous transcription model. Applications that require live captions, immediate partial transcripts, or interactive voice input should use a streaming speech-recognition model such as Qwen-Audio-3.1-ASR-Flash-Streaming.
Does Qwen-Audio-3.1-ASR-Flash-Filetrans support speaker separation and timestamps?
Yes. The model supports speaker separation and timestamps, along with multilingual recognition, regional Chinese dialect recognition, punctuation prediction, text normalization, hot-word enhancement, and context enhancement. Alibaba Cloud recommends keeping recordings to about two hours when speaker diarization is enabled.
What is Qwen-Audio-3.1-ASR-Flash-Filetrans used for?
Qwen-Audio-3.1-ASR-Flash-Filetrans is an asynchronous speech-to-text model for transcribing long audio and video recordings. It is suitable for meetings, lectures, interviews, customer calls, broadcasts, subtitles, and searchable audio archives.
What file types and size limits does Qwen-Audio-3.1-ASR-Flash-Filetrans support?
The model accepts publicly accessible HTTP or HTTPS URLs for audio and video files, including AAC, AMR, AVI, FLAC, FLV, M4A, MKV, MOV, MP3, MP4, MPEG, OGG, OPUS, WAV, WEBM, WMA, and WMV. The maximum documented file size is 2 GB, and the maximum duration is 12 hours.


Sources 5
Provider

About Qwen