Qwen3-LiveTranslate

Qwen3-LiveTranslate-Flash-Realtime

by Qwen · Legacy; still available; no longer recommended for new use

Alibaba Cloud's Qwen3-LiveTranslate-Flash-Realtime is a specialized legacy model for live multilingual interpretation. It processes streaming audio with optional image frames and returns translated text, synthesized speech, or both through a WebSocket API. The model supports 18 documented languages, a 53,248-token context window, and modality-based regional pricing, but it lacks tool use, structured outputs, fine-tuning, caching, and batch inference. New deployments should compare the newer Qwen3.5 and Qwen3.8 LiveTranslate models.

Text Speech Reasoning Coding
Qwen3-LiveTranslate-Flash-Realtime is a specialized translation model from Alibaba Cloud Model Studio for applications that need low-latency, continuous interpretation. It processes streaming audio and optional visual context, returning translated text, synthesized audio, or both. The model remains available, but Alibaba Cloud classifies it as legacy and no longer recommends it as the default choice for new projects.
Outputs

What Qwen3-LiveTranslate-Flash-Realtime can produce

Text Speech
Inputs

What it can understand

Images Audio Video Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family Qwen3-LiveTranslate
Model type Other
Context window 53K tokens
Maximum output 4K tokens
Release date 2025-09-23
Status Legacy; still available; no longer recommended for new use
Knowledge cutoff notes

Alibaba Cloud documents model context limits, supported modalities, pricing, and language coverage but does not publish a knowledge cutoff for this specialized real-time translation model.

Model notes

The canonical model ID qwen3-livetranslate-flash-realtime is currently equivalent to the qwen3-livetranslate-flash-realtime-2025-09-22 snapshot. Alibaba Cloud lists it as a legacy model that remains accessible but is no longer recommended for new use. The real-time API requires streaming and supports audio input with optional image frames; video context is supplied as image frames from a video stream. The model supports 18 documented languages, with audio output availability varying by language. Editorial scores are comparative estimates for this specialized translation model, not vendor benchmarks.

Cost

Model pricing

Input China (Beijing): CNY 64 per 1M input audio tokens and CNY 8 per 1M input image tokens. Singapore: CNY 73.392 per 1M input audio tokens and CNY 9.541 per 1M input image tokens.
Output China (Beijing): CNY 64 per 1M output text tokens and CNY 240 per 1M output audio tokens. Singapore: CNY 73.392 per 1M output text tokens and CNY 278.891 per 1M output audio tokens.
Model guide

Qwen3-LiveTranslate-Flash-Realtime: Legacy Model for Live Speech Translation

Qwen3-LiveTranslate-Flash-Realtime is Alibaba Cloud's legacy real-time audiovisual translation model. It accepts streaming audio and optional image frames, then returns translated text, synthesized speech, or both. Its WebSocket interface, approximately 53K-token context window, modality-based pricing, and support for 18 documented languages make it suitable for live interpretation, although Alibaba Cloud recommends newer Qwen3.5 and Qwen3.8 LiveTranslate models for new deployments.

What Qwen3-LiveTranslate-Flash-Realtime does

Qwen3-LiveTranslate-Flash-Realtime is designed for live audiovisual translation rather than general-purpose conversation. An application sends audio over a real-time connection, and the model produces translation results while the session is still in progress. This makes it appropriate for situations where waiting for a complete recording before translating would be impractical.

The model can return translated text, synthesized speech, or both. For example, a conference application could display translated subtitles while also playing the translation through a speaker. A voice-communication product could use the audio output for near-real-time interpretation, while a video workflow could use text output to create translated captions.

Alibaba Cloud provides the model through Model Studio's real-time WebSocket API. The canonical model identifier is qwen3-livetranslate-flash-realtime, which currently resolves to the qwen3-livetranslate-flash-realtime-2025-09-22 snapshot. The alias and dated snapshot should be treated as the same currently documented model version rather than two separate models.

Inputs, outputs, and visual context

Audio is the required input for translation. The model can also accept image frames as optional visual context. These frames may come from individual images or from a video stream sampled into images. Visual context can help the system interpret visible text, gestures, objects, or other scene information that affects the meaning of spoken content.

Video is therefore not described as a separate native video input in the supplied model specification. Instead, an application supplies video context as a sequence of image frames alongside the audio stream. The available output modalities are text and synthesized audio; the model does not generate images or video.

CapabilityDocumented support
Audio inputYes; required for the translation workflow
Image inputYes; optional visual context
Video inputVideo context can be supplied as image frames
Text inputNot listed as a standalone input modality in the model record
Text outputYes; translated text is streamed
Audio outputYes; synthesized translated speech is streamed
Image or video outputNo

Language coverage

Alibaba Cloud's documentation identifies 18 languages for this legacy real-time model: English, Chinese, Russian, French, German, Portuguese, Spanish, Italian, Indonesian, Korean, Japanese, Vietnamese, Thai, Arabic, Cantonese, Hindi, Greek, and Turkish.

Language support is not completely uniform across every translation direction. The documentation distinguishes between languages that support audio and text output and languages that may be limited to text output depending on the selected source and target languages. Applications that require spoken output should therefore verify the specific language pair before committing to a production design.

Context window and technical limits

The model has a documented context window of 53,248 tokens. The maximum input is 49,152 tokens and the maximum output is 4,096 tokens. These figures describe the model's token limits, but real-time applications also need to account for the duration and structure of an active audio session, event timing, and the amount of visual context being sent.

SpecificationDocumented value
Context window53,248 tokens
Maximum input49,152 tokens
Maximum output4,096 tokens
Connection typeReal-time WebSocket
StreamingRequired
Stable model snapshotqwen3-livetranslate-flash-realtime-2025-09-22

The real-time endpoint is not simply a conventional request that accepts a complete prompt and returns one finished answer. Clients send audio-buffer events and receive incremental translation and audio events. A usable integration must process partial results as they arrive, preserve event ordering, and close the session after the service indicates that processing has finished.

Pricing by modality

Alibaba Cloud prices this model according to the modality and direction of token usage. The following rates are stated per one million tokens and are separated by region.

UsageChina, BeijingSingapore
Input audioCNY 64CNY 73.392
Input imageCNY 8CNY 9.541
Output textCNY 64CNY 73.392
Output audioCNY 240CNY 278.891

Alibaba Cloud documents conversion rules for the legacy model: each second of input or output audio consumes 12.5 tokens, while image usage is calculated using 28-by-28-pixel units at 0.5 tokens per unit. The practical cost depends heavily on the amount of audio processed and whether synthesized speech is returned. Audio output is substantially more expensive per token than text output, so text-only subtitles or transcripts may be preferable when spoken translation is not required.

The rates above are provider-documented regional prices, not a recurring subscription price. Promotional offers, free quotas, account eligibility, and regional billing conditions may change.

Reasoning, coding, and tool support

This model is optimized for translation and audiovisual interpretation, not broad reasoning or software development. Its editorial reasoning and coding scores are both rated 1 because those capabilities are not central to its purpose; these are comparative editorial assessments, not Alibaba Cloud benchmark results.

According to the supplied capability documentation, the model does not support function calling, structured outputs, web search, fine-tuning, context caching, or batch inference. It should not be selected as the foundation for an agent that needs to call external tools, return strict machine-readable schemas, or perform offline batch processing.

  • Reasoning: Suitable for interpreting and translating streaming speech, but not positioned as a general reasoning model.
  • Coding: Not a coding model and not appropriate for software-generation workflows.
  • Function calling and tools: Not supported.
  • Structured output: Not supported according to the model capability record.
  • Fine-tuning, caching, and batch inference: Not supported.

Its main performance trade-off is specialization. A dedicated streaming translation model can be a better fit than a general conversational model when the application needs continuous audio translation and synthesized speech, but it offers fewer general-purpose controls and tools.

Strengths and trade-offs

The strongest feature of Qwen3-LiveTranslate-Flash-Realtime is its end-to-end real-time workflow. It combines streaming audio input, optional visual context, incremental text translation, and synthesized speech output in one specialized model. That combination reduces the need to assemble separate speech, translation, and speech-synthesis components for a basic live interpretation product.

The model also covers a relatively broad documented set of 18 languages and can use image frames when speech alone does not provide enough context. Its 53,248-token context window is substantial for a streaming model, although context size alone does not guarantee a particular translation quality or latency.

There are important trade-offs. The real-time API requires a WebSocket streaming integration, which is more operationally involved than sending a single ordinary completion request. Audio output is the most expensive priced modality. Language pairs may have different audio-output availability. The model also lacks the tools and structured-output features expected in broader application platforms.

Speed and cost scores in the supplied record are editorial estimates: speed 8 out of 10 and cost 4 out of 10. They should be read as comparative guidance for this specialized model, not as provider-published benchmarks. The high speed assessment reflects its real-time orientation, while the middling cost assessment reflects modality-based billing and the relatively high price of synthesized audio output.

When to choose this model

Choose Qwen3-LiveTranslate-Flash-Realtime when the central requirement is live multilingual speech interpretation and the application can work with a streaming WebSocket connection. Good fits include:

  • Live interpretation for conferences, meetings, classrooms, and events.
  • Voice communication between speakers of different languages.
  • Real-time subtitles and translated captions for streams or video calls.
  • Multilingual media applications that need translated speech as well as text.
  • Audiovisual translation where image frames provide useful scene or on-screen-text context.

It is less suitable for general-purpose chat, complex reasoning, coding assistants, tool-using agents, strict JSON workflows, fine-tuned applications, or large offline translation jobs. For those requirements, a general model or a translation pipeline with explicit batch and tool support may be more appropriate.

Where it fits in the current lineup

Qwen3-LiveTranslate-Flash-Realtime remains listed as available in Alibaba Cloud Model Studio, but it is classified as a legacy model and is no longer the recommended choice for new use cases. Alibaba Cloud identifies Qwen3.5-LiveTranslate-Flash-Realtime and Qwen3.8-LiveTranslate-Flash-Realtime as newer recommended alternatives.

The practical implication is that this model is most defensible when an existing application already depends on its behavior, pricing, supported language set, or snapshot. New projects should compare the newer LiveTranslate models before standardizing on the Qwen3 version. The supplied research does not provide enough detailed specifications for those newer models to make a numerical quality, price, or capability comparison here.

Bottom line

Qwen3-LiveTranslate-Flash-Realtime is a focused real-time translation model rather than a general AI assistant. It accepts streaming audio, can incorporate image frames, and returns translated text or synthesized speech through a WebSocket interface. Its documented limits and modality-specific pricing make it possible to plan a deployment, but audio output can materially increase cost and the model lacks tools, structured outputs, and batch processing.

For an existing live interpretation or audiovisual translation workflow, it remains a practical legacy option. For a new application, its status means the newer Qwen3.5 and Qwen3.8 LiveTranslate models deserve evaluation first.


Answers to Frequently Asked Questions

Should new projects use Qwen3-LiveTranslate-Flash-Realtime?
Qwen3-LiveTranslate-Flash-Realtime is classified as a legacy model and is no longer the recommended choice for new use cases. Existing applications may continue using it when they depend on its behavior, pricing, language coverage, or model snapshot, but new projects should first evaluate the newer Qwen3.5-LiveTranslate-Flash-Realtime and Qwen3.8-LiveTranslate-Flash-Realtime models.
How much does Qwen3-LiveTranslate-Flash-Realtime cost?
Pricing is based on token usage and modality. In China, Beijing, the documented rates per one million tokens are CNY 64 for input audio, CNY 8 for input images, CNY 64 for output text, and CNY 240 for output audio. In Singapore, the corresponding rates are CNY 73.392, CNY 9.541, CNY 73.392, and CNY 278.891. Synthesized audio is substantially more expensive than text output.
Which languages does Qwen3-LiveTranslate-Flash-Realtime support?
The documented language set includes English, Chinese, Russian, French, German, Portuguese, Spanish, Italian, Indonesian, Korean, Japanese, Vietnamese, Thai, Arabic, Cantonese, Hindi, Greek, and Turkish. Audio and text output availability can vary by language pair, so applications requiring spoken translation should verify the selected source and target languages.
What is Qwen3-LiveTranslate-Flash-Realtime used for?
Qwen3-LiveTranslate-Flash-Realtime is designed for live audiovisual translation. It accepts streaming audio over a real-time WebSocket connection and can return translated text, synthesized speech, or both for use cases such as live interpretation, translated subtitles, video calls, conferences, and multilingual media.
What input and output modalities does Qwen3-LiveTranslate-Flash-Realtime support?
Audio is the required input, while image frames can be supplied as optional visual context from individual images or sampled video. The model streams translated text and synthesized audio as outputs. It does not generate images or video, and video is provided indirectly as a sequence of image frames.


Sources 5
Provider

About Qwen