Qwen3-ASR

Qwen3-ASR-1.7B

by Qwen · Current; open-weight and downloadable

Qwen3-ASR-1.7B is a roughly 2-billion-parameter open-weight speech recognition model for multilingual transcription and language identification. It supports 30 languages, 22 Chinese dialects, offline and streaming inference, long audio, singing voice, and songs with background music. It is designed for self-hosted audio-to-text workflows, while timestamps require a separate forced-aligner model and no hosted API price is specified for this checkpoint.

Text Reasoning Coding
Qwen3-ASR-1.7B is a roughly 2-billion-parameter speech recognition checkpoint designed to convert audio into text. It is aimed at developers and organizations that need downloadable multilingual ASR for local processing, batch transcription, or real-time streaming. The model supports 30 languages and 22 Chinese dialects, but it is a specialized transcription system rather than a general-purpose conversational, reasoning, or text-generation model.
Outputs

What Qwen3-ASR-1.7B can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Fine-tuning Batch API
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen3-ASR
Model type Other
Context window tokens
Maximum output tokens
Release date 2026-01-29
Status Current; open-weight and downloadable
Knowledge cutoff notes

This is an automatic speech recognition model rather than a knowledge-grounded text assistant. The official model documentation does not publish a conventional knowledge cutoff date.

Model notes

Qwen3-ASR-1.7B is a specialized automatic speech recognition checkpoint with approximately 2B parameters and BF16 weights. It supports 30 languages and 22 Chinese dialects, offline and streaming inference, batch processing, long audio, speech, singing voice, and songs with background music. Word- and character-level timestamps require the separate Qwen3-ForcedAligner-0.6B model. The official repository documents local Transformers and vLLM deployment and provides an OpenAI-compatible local serving interface. The checkpoint is released under Apache 2.0. The documented max_new_tokens value of 4096 is an inference example rather than a verified hard model output limit, so max_output_tokens is left unknown. Alibaba Cloud hosted ASR services use separate model identifiers and pricing and should not be conflated with this downloadable checkpoint.

Cost

Model pricing

Input No official hosted API price identified for this exact open-weight checkpoint
Output No separate output charge for the self-hosted checkpoint
Model guide

Qwen3-ASR-1.7B: Open-Weight Multilingual Speech Recognition for Local and Streaming Transcription

Qwen3-ASR-1.7B is an open-weight automatic speech recognition model from Qwen and Alibaba for multilingual transcription and language identification. It supports 30 languages, 22 Chinese dialects, offline and streaming inference, long audio, singing voice, and songs with background music.

What Qwen3-ASR-1.7B is

Qwen3-ASR-1.7B is an open-weight automatic speech recognition (ASR) model provided by Qwen, Alibaba's model research organization. ASR systems take spoken or other audio content as input and produce a written transcription as output. This checkpoint is distributed through the Qwen organization on Hugging Face under the Apache 2.0 license, making it suitable for self-hosted deployments subject to that license.

The model is the larger model in the Qwen3-ASR family and contains approximately 2 billion parameters according to the supplied model information. Its purpose is focused: it recognizes and transcribes audio rather than acting as a general-purpose language model. It can also identify the language of an audio recording.

Qwen describes Qwen3-ASR-1.7B as suitable for difficult acoustic conditions and for audio that may include singing or background music. These are provider or model-positioning claims; actual results will depend on recording quality, speakers, accents, noise, audio content, and the deployment configuration.

Languages and audio supported

Qwen3-ASR-1.7B supports automatic speech recognition and language identification for 30 languages. The documented language list includes Chinese, English, Cantonese, Arabic, German, French, Spanish, Portuguese, Indonesian, Italian, Korean, Russian, Thai, Vietnamese, Japanese, Turkish, Hindi, Malay, Dutch, Swedish, Danish, Finnish, Polish, Czech, Filipino, Persian, Greek, Hungarian, Macedonian, and Romanian.

Its dialect coverage is particularly detailed for Chinese. The model supports 22 Chinese dialects and regional varieties, including Anhui, Dongbei, Fujian, Gansu, Guizhou, Hebei, Henan, Hubei, Hunan, Jiangxi, Ningxia, Shandong, Shaanxi, Shanxi, Sichuan, Tianjin, Yunnan, and Zhejiang. The documented coverage also includes Cantonese accents from Hong Kong and Guangdong, as well as Wu and Minnan.

The supported audio scenarios extend beyond ordinary conversational speech. The supplied documentation identifies speech, singing voice, and songs with background music as supported inputs. That makes the model potentially useful for interviews, meetings, lectures, call recordings, multilingual media, and some music-related transcription tasks. It should not be assumed that every language or audio type will perform equally well in every acoustic environment.

How to deploy the model

Because Qwen3-ASR-1.7B is an open-weight checkpoint, it can be deployed locally rather than accessed only through a provider-hosted API. The official Qwen3-ASR project documents use with the qwen-asr package, Transformers-compatible tooling, and vLLM. The repository also provides interfaces for local files, URLs, base64-encoded audio, NumPy audio arrays, batch inference, and asynchronous serving.

Offline inference is appropriate when an application processes completed recordings, such as archived calls, uploaded interviews, or lecture files. Streaming inference is intended for applications that need partial transcription while audio is being received. According to the supplied documentation, streaming is currently provided through the vLLM backend. A local OpenAI-compatible serving interface is also documented, but that interface should not be confused with a Qwen-hosted public API or with a separate commercial pricing plan for this checkpoint.

Actual deployment requirements depend on the selected runtime, audio format, batch size, quantization or precision choices, concurrency, and available memory. The supplied information identifies BF16 weights but does not provide a verified universal hardware requirement or a single guaranteed real-time speed figure.

Transcription features and timestamp limits

The model directly produces text transcriptions and can process long audio. However, the research does not specify a conventional fixed context window or a hard maximum output-token limit for this checkpoint. Long-audio behavior therefore depends on the runtime, generation settings, available memory, segmentation strategy, and characteristics of the recording. The documented max_new_tokens value of 4096 is an inference example, not a verified model-wide output limit.

Word-level and character-level timestamps are not intrinsic outputs of Qwen3-ASR-1.7B according to the supplied documentation. Applications that require forced alignment should use the separate Qwen3-ForcedAligner-0.6B model alongside the ASR model. This distinction matters for subtitles, karaoke-style alignment, searchable transcripts with precise timing, and applications that need each word mapped to a point in the audio.

Capabilities and boundaries

CapabilityAssessment
Primary inputAudio
Primary outputText transcription
Language coverage30 languages and 22 Chinese dialects
Inference modesOffline and streaming
Audio types describedSpeech, singing voice, and songs with background music
Native image, video, or audio generationNot supported
Tool or function callingNot supported as a native model capability
Native structured JSON outputNot documented
Reasoning and codingNot the model's purpose

Qwen3-ASR-1.7B accepts multimodal input in the narrow sense that it takes audio rather than text alone, but its output is text only. It does not generate speech, music, images, or video. It is therefore best evaluated as a specialized audio-to-text component, not as a multimodal assistant.

It also does not provide general reasoning or coding capabilities. A surrounding application could send the resulting transcript to another language model for summarization, question answering, extraction, or code generation, but those would be capabilities of the downstream system rather than of Qwen3-ASR-1.7B itself.

Strengths and practical trade-offs

The model's main strength is the combination of broad multilingual coverage, substantial Chinese dialect coverage, downloadable weights, and support for both batch and streaming workflows. Self-hosting can be valuable when audio should remain within an organization's infrastructure, when a service needs predictable control over deployment, or when processing large volumes makes a hosted per-minute service less attractive.

The 1.7B checkpoint is positioned as the higher-accuracy member of the Qwen3-ASR family. That positioning suggests a quality-oriented choice within the family, although the supplied research does not provide benchmark tables or independently verified accuracy figures. A larger ASR model may require more memory and compute than a smaller alternative, so the practical trade-off is likely to involve recognition quality, latency, concurrency, and infrastructure cost rather than a single universally best choice.

For high-throughput or latency-sensitive applications, a smaller speech-recognition model or a managed speech API may be easier to operate. Such alternatives can reduce infrastructure work and may offer predictable service-level behavior, but they can introduce usage charges, data-handling considerations, vendor dependence, or less control over model deployment. Qwen3-ASR-1.7B is most compelling when open weights, local execution, language coverage, or specialized audio handling matter more than turnkey hosting.

Pricing and license

No official hosted API price was identified for this exact open-weight checkpoint in the supplied research. There is no separate output charge when the checkpoint is downloaded and self-hosted, although users still incur the costs of hardware, cloud compute, storage, bandwidth, monitoring, and engineering support. Alibaba Cloud offers separate hosted speech-recognition services, but their model identifiers and prices should not automatically be attributed to Qwen3-ASR-1.7B.

The checkpoint is released under the Apache 2.0 license according to the supplied model information. Organizations should still review the license text and their own compliance requirements before deploying it commercially, particularly when processing sensitive recordings or redistributing modified software and weights.

When to choose Qwen3-ASR-1.7B

Choose Qwen3-ASR-1.7B when you need a self-hosted or downloadable ASR model with broad multilingual coverage, substantial Chinese dialect support, offline transcription, or streaming through the documented vLLM path. It is a reasonable candidate for meeting and interview transcription, call-center processing, lecture archives, multilingual media workflows, and applications that must keep audio under local control.

  • Choose it for: multilingual transcription, language identification, long recordings, local batch processing, streaming ASR, and audio containing singing or background music.
  • Choose a hosted speech service instead when: you want managed scaling, a published per-minute price, minimal infrastructure work, or a service-level agreement.
  • Choose a smaller model instead when: available hardware, latency, or concurrent request capacity is more important than the highest-positioned accuracy option within this model family.
  • Add a forced-aligner model when: your application requires word- or character-level timestamps rather than plain transcription.
  • Use another language model after transcription when: you need summarization, reasoning, question answering, structured extraction, coding, or conversational responses.

Bottom line

Qwen3-ASR-1.7B is a focused, open-weight audio-to-text model rather than an all-purpose AI assistant. Its most important differentiators are multilingual and Chinese dialect coverage, local deployment, offline and streaming inference, and support for challenging audio types described by Qwen. Its main limitations are the absence of native general reasoning, coding, tool use, media generation, and precise timestamps without an additional aligner. For teams willing to operate their own inference stack, it offers a flexible foundation for multilingual transcription without a documented hosted API charge for the checkpoint itself.


Answers to Frequently Asked Questions

Does Qwen3-ASR-1.7B provide word-level or character-level timestamps?
No. Word-level and character-level timestamps are not intrinsic outputs of Qwen3-ASR-1.7B according to the supplied documentation. Applications that require precise timing should use the separate Qwen3-ForcedAligner-0.6B model alongside the ASR model.
How much does Qwen3-ASR-1.7B cost, and what license does it use?
The checkpoint is distributed under the Apache 2.0 license and can be downloaded and self-hosted without a separate per-output charge. Users still pay for hardware, cloud compute, storage, bandwidth, monitoring, and engineering. No official hosted API price was identified for this exact checkpoint, and Alibaba Cloud hosted speech services should not automatically be treated as pricing for Qwen3-ASR-1.7B.
Can Qwen3-ASR-1.7B run locally and support streaming transcription?
Yes. Qwen3-ASR-1.7B is an open-weight checkpoint that can be deployed locally using the qwen-asr package, Transformers-compatible tooling, or vLLM. It supports offline transcription, while streaming inference is documented through the vLLM backend. Deployment requirements depend on hardware, runtime, precision, batch size, and concurrency.
What is Qwen3-ASR-1.7B used for?
Qwen3-ASR-1.7B is an open-weight automatic speech recognition model that converts audio into text and can identify the audio language. It is designed for multilingual transcription, including meetings, interviews, lectures, call recordings, and some audio containing singing or background music.
Which languages and dialects does Qwen3-ASR-1.7B support?
The model supports automatic speech recognition and language identification for 30 languages, including English, Chinese, Cantonese, Arabic, German, French, Spanish, Japanese, Korean, Hindi, Persian, and Romanian. It also supports 22 Chinese dialects and regional varieties, including Cantonese accents from Hong Kong and Guangdong, Wu, and Minnan.


Sources 4
Provider

About Qwen