What Qwen3-LiveTranslate-Flash-Realtime does
Qwen3-LiveTranslate-Flash-Realtime is designed for live audiovisual translation rather than general-purpose conversation. An application sends audio over a real-time connection, and the model produces translation results while the session is still in progress. This makes it appropriate for situations where waiting for a complete recording before translating would be impractical.
The model can return translated text, synthesized speech, or both. For example, a conference application could display translated subtitles while also playing the translation through a speaker. A voice-communication product could use the audio output for near-real-time interpretation, while a video workflow could use text output to create translated captions.
Alibaba Cloud provides the model through Model Studio's real-time WebSocket API. The canonical model identifier is qwen3-livetranslate-flash-realtime, which currently resolves to the qwen3-livetranslate-flash-realtime-2025-09-22 snapshot. The alias and dated snapshot should be treated as the same currently documented model version rather than two separate models.
Inputs, outputs, and visual context
Audio is the required input for translation. The model can also accept image frames as optional visual context. These frames may come from individual images or from a video stream sampled into images. Visual context can help the system interpret visible text, gestures, objects, or other scene information that affects the meaning of spoken content.
Video is therefore not described as a separate native video input in the supplied model specification. Instead, an application supplies video context as a sequence of image frames alongside the audio stream. The available output modalities are text and synthesized audio; the model does not generate images or video.
| Capability | Documented support |
|---|---|
| Audio input | Yes; required for the translation workflow |
| Image input | Yes; optional visual context |
| Video input | Video context can be supplied as image frames |
| Text input | Not listed as a standalone input modality in the model record |
| Text output | Yes; translated text is streamed |
| Audio output | Yes; synthesized translated speech is streamed |
| Image or video output | No |
Language coverage
Alibaba Cloud's documentation identifies 18 languages for this legacy real-time model: English, Chinese, Russian, French, German, Portuguese, Spanish, Italian, Indonesian, Korean, Japanese, Vietnamese, Thai, Arabic, Cantonese, Hindi, Greek, and Turkish.
Language support is not completely uniform across every translation direction. The documentation distinguishes between languages that support audio and text output and languages that may be limited to text output depending on the selected source and target languages. Applications that require spoken output should therefore verify the specific language pair before committing to a production design.
Context window and technical limits
The model has a documented context window of 53,248 tokens. The maximum input is 49,152 tokens and the maximum output is 4,096 tokens. These figures describe the model's token limits, but real-time applications also need to account for the duration and structure of an active audio session, event timing, and the amount of visual context being sent.
| Specification | Documented value |
|---|---|
| Context window | 53,248 tokens |
| Maximum input | 49,152 tokens |
| Maximum output | 4,096 tokens |
| Connection type | Real-time WebSocket |
| Streaming | Required |
| Stable model snapshot | qwen3-livetranslate-flash-realtime-2025-09-22 |
The real-time endpoint is not simply a conventional request that accepts a complete prompt and returns one finished answer. Clients send audio-buffer events and receive incremental translation and audio events. A usable integration must process partial results as they arrive, preserve event ordering, and close the session after the service indicates that processing has finished.
Pricing by modality
Alibaba Cloud prices this model according to the modality and direction of token usage. The following rates are stated per one million tokens and are separated by region.
| Usage | China, Beijing | Singapore |
|---|---|---|
| Input audio | CNY 64 | CNY 73.392 |
| Input image | CNY 8 | CNY 9.541 |
| Output text | CNY 64 | CNY 73.392 |
| Output audio | CNY 240 | CNY 278.891 |
Alibaba Cloud documents conversion rules for the legacy model: each second of input or output audio consumes 12.5 tokens, while image usage is calculated using 28-by-28-pixel units at 0.5 tokens per unit. The practical cost depends heavily on the amount of audio processed and whether synthesized speech is returned. Audio output is substantially more expensive per token than text output, so text-only subtitles or transcripts may be preferable when spoken translation is not required.
The rates above are provider-documented regional prices, not a recurring subscription price. Promotional offers, free quotas, account eligibility, and regional billing conditions may change.
Reasoning, coding, and tool support
This model is optimized for translation and audiovisual interpretation, not broad reasoning or software development. Its editorial reasoning and coding scores are both rated 1 because those capabilities are not central to its purpose; these are comparative editorial assessments, not Alibaba Cloud benchmark results.
According to the supplied capability documentation, the model does not support function calling, structured outputs, web search, fine-tuning, context caching, or batch inference. It should not be selected as the foundation for an agent that needs to call external tools, return strict machine-readable schemas, or perform offline batch processing.
- Reasoning: Suitable for interpreting and translating streaming speech, but not positioned as a general reasoning model.
- Coding: Not a coding model and not appropriate for software-generation workflows.
- Function calling and tools: Not supported.
- Structured output: Not supported according to the model capability record.
- Fine-tuning, caching, and batch inference: Not supported.
Its main performance trade-off is specialization. A dedicated streaming translation model can be a better fit than a general conversational model when the application needs continuous audio translation and synthesized speech, but it offers fewer general-purpose controls and tools.
Strengths and trade-offs
The strongest feature of Qwen3-LiveTranslate-Flash-Realtime is its end-to-end real-time workflow. It combines streaming audio input, optional visual context, incremental text translation, and synthesized speech output in one specialized model. That combination reduces the need to assemble separate speech, translation, and speech-synthesis components for a basic live interpretation product.
The model also covers a relatively broad documented set of 18 languages and can use image frames when speech alone does not provide enough context. Its 53,248-token context window is substantial for a streaming model, although context size alone does not guarantee a particular translation quality or latency.
There are important trade-offs. The real-time API requires a WebSocket streaming integration, which is more operationally involved than sending a single ordinary completion request. Audio output is the most expensive priced modality. Language pairs may have different audio-output availability. The model also lacks the tools and structured-output features expected in broader application platforms.
Speed and cost scores in the supplied record are editorial estimates: speed 8 out of 10 and cost 4 out of 10. They should be read as comparative guidance for this specialized model, not as provider-published benchmarks. The high speed assessment reflects its real-time orientation, while the middling cost assessment reflects modality-based billing and the relatively high price of synthesized audio output.
When to choose this model
Choose Qwen3-LiveTranslate-Flash-Realtime when the central requirement is live multilingual speech interpretation and the application can work with a streaming WebSocket connection. Good fits include:
- Live interpretation for conferences, meetings, classrooms, and events.
- Voice communication between speakers of different languages.
- Real-time subtitles and translated captions for streams or video calls.
- Multilingual media applications that need translated speech as well as text.
- Audiovisual translation where image frames provide useful scene or on-screen-text context.
It is less suitable for general-purpose chat, complex reasoning, coding assistants, tool-using agents, strict JSON workflows, fine-tuned applications, or large offline translation jobs. For those requirements, a general model or a translation pipeline with explicit batch and tool support may be more appropriate.
Where it fits in the current lineup
Qwen3-LiveTranslate-Flash-Realtime remains listed as available in Alibaba Cloud Model Studio, but it is classified as a legacy model and is no longer the recommended choice for new use cases. Alibaba Cloud identifies Qwen3.5-LiveTranslate-Flash-Realtime and Qwen3.8-LiveTranslate-Flash-Realtime as newer recommended alternatives.
The practical implication is that this model is most defensible when an existing application already depends on its behavior, pricing, supported language set, or snapshot. New projects should compare the newer LiveTranslate models before standardizing on the Qwen3 version. The supplied research does not provide enough detailed specifications for those newer models to make a numerical quality, price, or capability comparison here.
Bottom line
Qwen3-LiveTranslate-Flash-Realtime is a focused real-time translation model rather than a general AI assistant. It accepts streaming audio, can incorporate image frames, and returns translated text or synthesized speech through a WebSocket interface. Its documented limits and modality-specific pricing make it possible to plan a deployment, but audio output can materially increase cost and the model lacks tools, structured outputs, and batch processing.
For an existing live interpretation or audiovisual translation workflow, it remains a practical legacy option. For a new application, its status means the newer Qwen3.5 and Qwen3.8 LiveTranslate models deserve evaluation first.

