Qwen3-LiveTranslate

Qwen3-LiveTranslate-Flash

by Qwen · Current stable model

A specialized Alibaba Cloud model for translating recorded or uploaded audio and video. Qwen3-LiveTranslate-Flash supports 18 languages, visual context for video, streaming translated text, and synthesized audio for supported target languages. It uses a required OpenAI-compatible streaming API, has a 53,248-token context length and 4,096-token maximum output, and uses regional token-based Model Studio billing rather than one verified universal monetary rate.

Text Speech Reasoning Coding
Qwen3-LiveTranslate-Flash is a focused translation model in Alibaba Cloud Model Studio. It accepts an audio or video file, processes the spoken content through a required streaming API, and returns translated text, translated audio, or both when the selected language supports those outputs. Its main distinction is that it combines multilingual speech translation with optional video understanding, while remaining separate from the Qwen3-LiveTranslate-Flash-Realtime model intended for simultaneous interpretation.
Outputs

What Qwen3-LiveTranslate-Flash can produce

Text Speech
Inputs

What it can understand

Audio Video Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
Specifications

Technical details

Model family Qwen3-LiveTranslate
Model type Other
Context window 53K tokens
Maximum output 4K tokens
Release date 2025-09-22
Status Current stable model
Knowledge cutoff notes

Alibaba Cloud's official documentation does not state a knowledge cutoff for this specialized audio and video translation model.

Model notes

The canonical API model identifier is qwen3-livetranslate-flash. Alibaba Cloud documents qwen3-livetranslate-flash-2025-12-01 as a snapshot with the same documented capabilities. The model uses an OpenAI-compatible chat completions endpoint, requires streaming, and does not support the DashScope interface. It accepts one user message containing audio or video. It supports translation across 18 languages; some target languages provide text output only. Video translation can use visual context to improve interpretation of ambiguous terms. The model is distinct from qwen3-livetranslate-flash-realtime, which is intended for live simultaneous interpretation over a realtime interface.

Cost

Model pricing

Input Audio input and output are billed at 12.5 tokens per second, with audio shorter than one second billed as one second. Video usage additionally consumes video tokens based on sampled frames and resolution. Monetary rates depend on the applicable Alibaba Cl
Output Audio output is billed at 12.5 tokens per second. Text output uses the applicable Model Studio token rate when charged separately. No standalone fixed monetary rate was verified for this exact model in the model documentation.
Model guide

Qwen3-LiveTranslate-Flash for Streaming Audio and Video Translation

Qwen3-LiveTranslate-Flash is Alibaba Cloud's specialized Qwen3 model for translating uploaded or recorded audio and video into streaming translated text, synthesized speech, or both. It supports 18 languages, can use visual context from video, and is designed for file-translation workflows rather than general chat or unrestricted live interpretation.

What Qwen3-LiveTranslate-Flash does

Qwen3-LiveTranslate-Flash is a specialized audio and video translation model provided through Alibaba Cloud Model Studio. Its purpose is narrow and practical: take recorded or uploaded audiovisual material and translate the spoken content into another language.

The model can return streaming translated text, synthesized translated speech, or both. For example, an application could use it to create translated subtitles from a recorded interview, generate a translated voice track for a video, or display translated dialogue as it is produced. The provider documents support for 18 languages, although output options vary by target language and some languages provide text output only.

This is not a general-purpose conversational model. It is also not primarily a speech-recognition model for producing transcripts without translation. Its documented role is translation, making it more appropriate for multilingual media processing than for open-ended question answering, coding, or image generation.

Where it fits in the Qwen lineup

Qwen3-LiveTranslate-Flash belongs to the Qwen3-LiveTranslate family and is the file-oriented translation option in that family. Alibaba Cloud separately documents qwen3-livetranslate-flash-realtime for real-time simultaneous interpretation over a realtime interface. That distinction matters: the current model is intended for audio and video files, while the realtime sibling is the more relevant choice for live conversations or ongoing interpretation.

The canonical API identifier for this model is qwen3-livetranslate-flash. Alibaba Cloud also documents the snapshot identifier qwen3-livetranslate-flash-2025-12-01, described as having the same documented capabilities. The model was released on September 22, 2025, according to the supplied Qwen release information.

Supported inputs and outputs

CapabilityDocumented behavior
Audio inputSupported
Video inputSupported
Text inputNot listed as a supported input modality for this model
Translated text outputSupported through streaming responses
Translated audio outputSupported for languages with speech-output availability
Image outputNot supported
Video outputNot supported
StreamingRequired by the API

A request contains one user message with audio or video content. Video translation can use visual context in addition to the soundtrack. This may help when the visual scene clarifies an ambiguous term, speaker reference, or other meaning that cannot be determined from audio alone. The model still produces translation rather than an edited or newly rendered video file, so an application must handle subtitle placement, audio-video synchronization, or final media production separately.

Audio output is synthesized speech rather than a finished dubbed video. If a target language supports only text output, the application receives translated text without a corresponding generated voice track.

Context, output limits, and API behavior

The documented context length is 53,248 tokens, and the maximum output is 4,096 tokens. These are model-token limits rather than direct measures of minutes of audio or video. The amount of media that fits in a request depends on factors such as duration, sampling, resolution, and the way the service converts audiovisual content into tokens.

Audio shorter than one second is billed as one second. Video usage additionally consumes video tokens based on sampled frames and resolution. The provider requires streaming for this model, so clients should be prepared to read incremental output rather than expecting a conventional single, non-streaming response.

The API is OpenAI-compatible and uses a chat-completions-style interface, which can reduce integration work for applications that already understand that request pattern. However, the model documentation states that it does not support the DashScope interface. OpenAI compatibility should therefore not be interpreted as proof that every OpenAI SDK feature or every Model Studio interface is available; the model-specific API requirements still apply.

Pricing and cost

No single fixed monetary price for Qwen3-LiveTranslate-Flash was verified in the supplied model documentation. Alibaba Cloud describes billing in token terms: audio input and audio output are each charged at 12.5 tokens per second, with audio shorter than one second billed as one second. Video requests also consume video tokens according to sampled frames and resolution.

The final monetary charge depends on the applicable Alibaba Cloud Model Studio region and pricing schedule. Text output may use the relevant Model Studio token rate when it is charged separately. Because the supplied sources do not establish one universal currency price for this exact model, quoting a per-million-token or per-minute amount would risk mixing regional schedules or inventing a rate.

In practical terms, audio-only translation should generally be easier to estimate than video translation because video adds frame-based usage. Applications processing large video libraries should account for both spoken-audio duration and visual sampling rather than budgeting only for speech seconds.

Capabilities and trade-offs

The model's strongest capability is specialization. Instead of asking a general language model to interpret an audio or video attachment, a translation-focused model is designed around the specific workflow of converting spoken content into another language. Support for translated speech makes it useful when subtitles alone are insufficient, while visual context can assist with video passages whose meaning depends partly on what appears on screen.

Its speed rating in the supplied research is an editorial assessment of 8 out of 10, not a provider-published benchmark. That assessment indicates a model positioned for responsive streaming workloads, but actual latency will depend on media length, output type, network conditions, region, and service load. The research does not provide a verified cost score.

The model has a reasoning score of 2 out of 10 and a coding score of 1 out of 10 in the supplied editorial data. These are not official benchmark results and should not be read as claims that the model cannot handle any reasoning or code-related text. They indicate that general reasoning and coding are outside its intended specialization. For translation pipelines, the relevant evaluation questions are language coverage, fidelity, latency, audio quality, and handling of the application's media formats.

Tool and function support is not documented for this model in the supplied research. The model should therefore be treated as a translation endpoint rather than an agent that independently browses the web, calls external tools, edits files, or invokes business functions.

Best use cases

  • Translated subtitles: Convert recorded lectures, interviews, presentations, or other spoken media into streaming translated text.
  • Translated voice tracks: Generate synthesized speech for supported target languages as part of a separate dubbing or localization pipeline.
  • Multilingual media archives: Process stored audio and video so users can search, review, or consume content in another language.
  • Video translation with visual context: Translate material where on-screen context can help resolve ambiguous speech or references.
  • Streaming application interfaces: Show partial translated results as they arrive instead of waiting for the complete file to finish.

For production systems, developers should test representative recordings rather than relying only on the language count. Accent, background noise, overlapping speakers, specialist vocabulary, speech rate, and the intended target-language output mode can all affect whether the result is suitable without human review.

When to choose this model

Choose Qwen3-LiveTranslate-Flash when the central task is translating recorded or uploaded audio and video, especially when the application benefits from streaming text or synthesized speech. It is a sensible fit for subtitle generation, multilingual media localization, and translation services that need video context.

Choose the separate Qwen3-LiveTranslate-Flash-Realtime model when the requirement is live, simultaneous interpretation rather than processing an audio or video file. The supplied research specifically distinguishes that model as the realtime option.

Choose a general-purpose language model when translation is only one step in a broader workflow involving substantial reasoning, coding, document transformation, or tool use. Likewise, a dedicated speech-recognition system may be more appropriate when the goal is accurate transcription in the original language without translation. Qwen3-LiveTranslate-Flash should not be selected for image generation, video generation, general chat, or complex software development because those are outside its documented purpose.

Limitations to plan for

  • The API requires streaming, which adds implementation requirements for clients that expect one complete response.
  • Some target languages provide translated text only and do not provide synthesized speech output.
  • Video billing and processing depend on sampled frames and resolution, so cost and token use are less predictable than audio-only processing.
  • No universal fixed monetary price was verified; regional Model Studio pricing must be checked before deployment.
  • The model is not documented as supporting tools, function calling, general text input, or unrestricted realtime interpretation.
  • The 53,248-token context length and 4,096-token maximum output do not directly guarantee a particular media duration or translation quality.

Overall, Qwen3-LiveTranslate-Flash is best understood as a fast, specialized translation component for audiovisual files. Its value comes from combining audio and video input, streaming translated results, and optional speech output—not from broad reasoning or general-purpose assistant behavior.


Answers to Frequently Asked Questions

What are the main limitations of Qwen3-LiveTranslate-Flash?
The API requires streaming, some languages do not support synthesized speech, and video usage is affected by frame sampling and resolution. The model is intended for translation rather than general chat, coding, tool use, unrestricted realtime interpretation, or standalone speech recognition.
How is Qwen3-LiveTranslate-Flash priced?
Alibaba Cloud bills the model in token terms. Audio input and audio output are each charged at 12.5 tokens per second, with audio shorter than one second billed as one second. Video also consumes tokens based on sampled frames and resolution, while the final monetary price depends on the applicable Model Studio region and pricing schedule.
What inputs and outputs does Qwen3-LiveTranslate-Flash support?
The model supports audio and video input, streaming translated text output, and synthesized translated audio for languages with speech-output support. It does not produce edited video or image output, and some target languages offer text translation only.
What is Qwen3-LiveTranslate-Flash used for?
Qwen3-LiveTranslate-Flash is a specialized Alibaba Cloud model for translating recorded or uploaded audio and video into another language. It can produce streaming translated text, synthesized translated speech, or both, depending on the target language.
Does Qwen3-LiveTranslate-Flash support real-time interpretation?
Qwen3-LiveTranslate-Flash is designed for audio and video files and requires streaming responses. For live conversations or simultaneous interpretation, Alibaba Cloud documents the separate qwen3-livetranslate-flash-realtime model.


Sources 4
Provider

About Qwen