Qwen3-ASR

Qwen3-ASR-0.6B

by Qwen · Current, open-weight, Apache 2.0 licensed

Compact open-weight Qwen speech recognition model supporting multilingual transcription, language identification, offline and streaming inference, and challenging audio such as singing and music-backed speech.

Text Reasoning Coding
Qwen3-ASR-0.6B is a 0.6-billion-parameter speech recognition model released by the Qwen team at Alibaba under the Apache 2.0 license. It supports 30 languages and 22 Chinese dialects, unified offline and streaming inference, long-audio transcription, and challenging audio such as singing or speech over background music. Its compact size makes it especially relevant for self-hosted, cost-sensitive, and high-throughput speech-to-text deployments.
Outputs

What Qwen3-ASR-0.6B can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Batch API
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family Qwen3-ASR
Model type Other
Context window 66K tokens
Maximum output tokens
Release date 2026-01-29
Status Current, open-weight, Apache 2.0 licensed
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this speech recognition checkpoint. Its speech-recognition training and supported-language documentation should not be interpreted as a conventional LLM knowledge cutoff.

Model notes

Qwen3-ASR-0.6B is the compact member of the Qwen3-ASR family and is developed by the Qwen team at Alibaba. The canonical open-weight checkpoint is Qwen/Qwen3-ASR-0.6B; Qwen/Qwen3-ASR-0.6B-hf is a Transformers-native compatible checkpoint and should not be treated as a separate underlying model. The model supports automatic language identification and speech recognition for 30 languages plus 22 Chinese dialects, with support for speech, singing voice, and songs with background music. It supports unified offline and streaming inference and long-audio transcription. Word-level or segment timestamps require the separate Qwen3-ForcedAligner-0.6B model. No official hosted per-token or per-minute price was identified for this open-weight checkpoint; deployment is generally self-hosted or through third-party infrastructure. The 65536 value is the configured text-model maximum position length, not a guaranteed maximum audio duration. Official examples use configurable generation limits such as 256 or 1024 tokens rather than documenting a fixed model-wide maximum output-token limit. Editorial scores are comparative estimates for ASR use and are not vendor specifications.

Model guide

Qwen3-ASR-0.6B: Compact Open-Weight Speech Recognition for Multilingual and Streaming Transcription

Qwen3-ASR-0.6B is Alibaba Qwen’s compact open-weight automatic speech recognition model for multilingual transcription, language identification, offline processing, long-audio transcription, and streaming speech recognition.

What is Qwen3-ASR-0.6B?

Qwen3-ASR-0.6B is an automatic speech recognition (ASR) model from the Qwen team at Alibaba. ASR models convert spoken audio into text and may also identify the language being spoken. This model is designed specifically for speech recognition rather than general-purpose conversation, coding, image analysis, or content generation.

The model has 0.6 billion parameters and is distributed as an open-weight checkpoint under the Apache 2.0 license. The canonical checkpoint is Qwen/Qwen3-ASR-0.6B. A Transformers-compatible checkpoint named Qwen/Qwen3-ASR-0.6B-hf is a compatible representation of the same underlying model, not a separate model in the family.

Within the Qwen3-ASR lineup, the 0.6B version is the compact member. Its positioning favors deployment efficiency, throughput, and operating cost over the broader capabilities expected from a general-purpose language model or a larger speech model.

Languages and audio the model can process

According to the supplied model documentation, Qwen3-ASR-0.6B supports automatic language identification and speech recognition for 30 languages plus 22 Chinese dialects. This makes it suitable for multilingual transcription systems that need to process recordings without requiring the language to be manually selected in advance.

The documented audio coverage extends beyond clean conversational speech. The model supports speech, singing voice, and songs or speech accompanied by background music. That does not mean every recording will transcribe equally well: noisy environments, overlapping speakers, strong accents, reverberation, and poor microphones can still reduce accuracy. However, the supported audio categories make the model more applicable to media, call, meeting, and user-generated audio workflows than a system restricted to clear speech.

Offline, streaming, and long-audio transcription

Qwen3-ASR-0.6B supports unified offline and streaming inference. Offline inference processes an available recording and returns a transcription, which is useful for files, archives, batch jobs, and recorded meetings. Streaming inference processes audio progressively, making it appropriate for live captions, voice interfaces, monitoring, and applications that need partial results while someone is speaking.

The model documentation also describes support for long-audio transcription. The supplied configuration reports a maximum text position length of 65,536, but this value should not be interpreted as a guaranteed maximum audio duration. Audio length depends on the preprocessing and inference pipeline, and the configured text-model position limit is not a direct promise that every 65,536-token-equivalent audio input can be processed in one operation.

Word-level or segment-level timestamps are not presented as a built-in feature of this checkpoint. The Qwen documentation identifies the separate Qwen3-ForcedAligner-0.6B model for forced alignment and timestamp generation. If precise timing is required for subtitles, searchable media, or karaoke-style applications, Qwen3-ASR-0.6B may therefore need to be combined with that separate alignment component.

Input, output, and capability profile

CapabilityQwen3-ASR-0.6B
Primary inputAudio
Primary outputText transcription
Audio inputSupported
Text inputSupported as part of the model interface or processing configuration
Image and video inputNot supported
Direct audio outputNot supported
StreamingSupported
Tool or function callingNot supported
Structured-output modeNo verified dedicated JSON mode
Fine-tuningNo verified support in the supplied research

The model’s output is text, not synthesized speech. It should not be selected when the application needs a spoken response, voice cloning, music generation, or another form of native audio generation. Its multimodal capability is limited to accepting audio for speech recognition; it is not a general image-and-audio reasoning model.

Context and output limits

The supplied configuration reports a 65,536-position text-model limit. This is the clearest published context-related value available for the checkpoint, but it should be treated carefully because Qwen3-ASR-0.6B is an audio-recognition model. The number does not establish a fixed maximum duration for an audio file.

No fixed, model-wide maximum output-token limit was identified in the supplied research. Official examples use configurable generation limits such as 256 or 1,024 tokens. Those values are runtime settings in example configurations rather than a guaranteed universal output ceiling. Applications should therefore size generation limits according to the expected transcription length and their chosen inference pipeline.

Speed, cost, and reasoning trade-offs

The 0.6B parameter count makes this the efficiency-oriented Qwen3-ASR option. The research rates its relative speed and cost favorably for ASR workloads, but those ratings are editorial estimates rather than vendor-published benchmarks. Actual throughput and latency depend on hardware, audio format, batching, quantization, concurrency, and whether inference is offline or streaming.

There is no official hosted per-token or per-minute price identified for the open-weight checkpoint. The model is generally intended for self-hosted deployment or use through third-party infrastructure, where costs depend on compute, storage, engineering, and the provider’s own pricing. It should not be advertised as having a specific free or paid API rate based solely on the existence of the public checkpoint.

Qwen3-ASR-0.6B is not a reasoning model in the usual language-model sense. Its job is to recognize and transcribe speech, not to solve multi-step problems, write software, browse the web, or call external tools. The supplied editorial scores rate its reasoning and coding usefulness low, which is appropriate for a specialized ASR system rather than a defect in its intended design.

Best use cases

  • Multilingual transcription: Convert recordings in supported languages and Chinese dialects into text.
  • Live captions: Use streaming inference for applications that need incremental speech recognition.
  • Offline archives: Process recorded meetings, interviews, lectures, media files, or other stored audio.
  • Language identification: Detect the spoken language as part of an ingestion or routing pipeline.
  • High-throughput deployments: Run a compact open-weight checkpoint on infrastructure controlled by the deploying organization.
  • Difficult audio categories: Experiment with singing voice and speech or songs containing background music.

Because it is open-weight and Apache 2.0 licensed, it may be a practical choice for teams that need more control over deployment or data handling than a hosted transcription API provides. The exact legal and operational suitability still depends on the application, infrastructure, and organization’s review of the license and source materials.

When to choose Qwen3-ASR-0.6B

Choose Qwen3-ASR-0.6B when the central requirement is speech-to-text and you value a compact open-weight model, multilingual coverage, streaming support, offline processing, and control over deployment. It is particularly well suited to teams that can operate their own inference stack or select a third-party host and want to avoid assuming a fixed vendor API price that is not published for the checkpoint.

Another option may be more appropriate when the application needs general-purpose reasoning, coding, image understanding, web search, tool use, or native speech generation. A separate alignment model or processing stage is also more suitable when reliable word-level or segment-level timestamps are essential. For production selection, test representative audio from the target languages, dialects, microphones, noise conditions, and music environments rather than relying only on the model’s supported-language list.

Bottom line

Qwen3-ASR-0.6B is a focused, compact speech recognition model rather than a general AI assistant. Its main value lies in multilingual transcription, language identification, offline and streaming inference, and support for audio that may include singing or background music. The open-weight Apache 2.0 release enables self-hosted deployment, while the absence of an official hosted price means operational cost must be calculated from the chosen infrastructure. Its 65,536 configured text position limit should not be mistaken for a fixed audio-duration limit, and timestamp extraction requires a separate forced-alignment component.


Answers to Frequently Asked Questions

Is Qwen3-ASR-0.6B free to use, and what license does it have?
Qwen3-ASR-0.6B is an open-weight checkpoint distributed under the Apache 2.0 license. No official hosted per-token or per-minute price was identified for the checkpoint, so deployment costs depend on hardware, storage, engineering, and any third-party hosting provider used.
Does Qwen3-ASR-0.6B provide timestamps or synthesized audio?
Qwen3-ASR-0.6B primarily outputs text transcription and does not natively provide synthesized audio. Word-level or segment-level timestamps are not presented as a built-in feature; applications requiring precise timing may need the separate Qwen3-ForcedAligner-0.6B model.
Does Qwen3-ASR-0.6B support streaming and long-audio transcription?
Yes. Qwen3-ASR-0.6B supports both offline and streaming inference, making it suitable for recorded files as well as live captions and voice interfaces. Its documented 65,536-position text limit should not be treated as a guaranteed maximum audio duration because the actual limit depends on preprocessing and the inference pipeline.
Which languages and audio types does Qwen3-ASR-0.6B support?
According to its documentation, Qwen3-ASR-0.6B supports automatic language identification and speech recognition for 30 languages plus 22 Chinese dialects. It can process conversational speech, singing voice, and speech or songs with background music, although noise, overlapping speakers, accents, reverberation, and poor microphones may reduce accuracy.
What is Qwen3-ASR-0.6B used for?
Qwen3-ASR-0.6B is a compact automatic speech recognition model designed to convert spoken audio into text and identify the spoken language. It supports multilingual transcription, offline processing, streaming inference, live captions, meeting and archive transcription, and audio containing singing or background music.


Sources 4
Provider

About Qwen