Qwen3.8-LiveTranslate

Qwen3.8-LiveTranslate-Flash-Realtime

by Qwen · Current stable model

A specialized Alibaba Cloud real-time translation model for streaming audio and image input, with text and speech output, 60-language understanding, 29 speech-output languages, and approximately 2.3-second provider-reported latency.

Text Speech Reasoning Coding
Qwen3.8-LiveTranslate-Flash-Realtime is built for live multilingual translation rather than ordinary chatbot conversations. It processes streaming audio and visual input, translates between supported languages, and can return either translated text or synthesized speech. Alibaba Cloud documents support for 60 languages, speech output in 29 languages, and latency as low as approximately 2.3 seconds. The model is best suited to meetings, live interpretation, translated voice communication, and other applications where response speed matters more than broad reasoning or tool capabilities.
Outputs

What Qwen3.8-LiveTranslate-Flash-Realtime can produce

Text Speech
Inputs

What it can understand

Images Audio Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

4/10 Reasoning
2/10 Coding
9/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.8-LiveTranslate
Model type Multimodal
Context window 53K tokens
Maximum output 4K tokens
Status Current stable model
Knowledge cutoff notes

Alibaba Cloud's official model documentation does not publish a knowledge cutoff for this real-time translation model. Its translation behavior is based on streaming input and configured source and target languages rather than a documented static knowledge cutoff.

Model notes

The exact canonical model ID is qwen3.8-livetranslate-flash-realtime. It accepts audio and image input and produces text and audio output. Alibaba Cloud documentation states that it understands 60 languages and supports speech output in 29 languages, with latency as low as approximately 2.3 seconds. Translation is accessed through the WebSocket Realtime API, with translated text streamed through response.text.delta or response.audio_transcript.delta and translated audio through response.audio.delta. The official model page lists function calling, structured outputs, web search, context caching, batch inference, and fine-tuning as unsupported. Pricing varies by region and modality. The model is distinct from qwen3.8-omni-flash-realtime, which is a broader real-time audio/video conversation model rather than the dedicated LiveTranslate model.

Cost

Model pricing

Input Singapore: audio input $7.50 per 1 million tokens; image input $0.55 per 1 million tokens. China (Beijing): audio input $5.653 per 1 million tokens; image input $0.466 per 1 million tokens.
Output Singapore: text output $20 per 1 million tokens; audio output $30 per 1 million tokens. China (Beijing): text output $14.133 per 1 million tokens; audio output $22.613 per 1 million tokens.
Model guide

Qwen3.8-LiveTranslate-Flash-Realtime: Low-Latency Audio and Visual Translation

Qwen3.8-LiveTranslate-Flash-Realtime is Alibaba Cloud Model Studio’s dedicated real-time translation model. It accepts streaming audio and image input, supports 60 languages, and returns translated text or speech through the WebSocket Realtime API. Its main advantage is low-latency simultaneous interpretation rather than general-purpose conversation, coding, or tool use.

What Qwen3.8-LiveTranslate-Flash-Realtime does

Qwen3.8-LiveTranslate-Flash-Realtime is a specialized real-time translation model provided through Alibaba Cloud Model Studio. Its purpose is to translate live audio and visual content as it arrives, rather than wait for a complete recording or document to be processed. The model is accessed through Alibaba Cloud’s WebSocket Realtime API, a persistent connection designed for streaming input and output.

In practical terms, an application can send audio or images to the model and receive translated text as a stream. It can also receive translated speech when audio output is selected. This makes the model suitable for live meetings, multilingual conversations, remote interpretation, and audiovisual content where waiting for a full batch translation would be inconvenient.

The model’s canonical identifier is qwen3.8-livetranslate-flash-realtime. It is listed as a current stable model in the supplied Alibaba Cloud Model Studio research. It belongs to the Qwen3.8-LiveTranslate family and is distinct from Qwen3.8-Omni-Flash-Realtime, which is positioned as a broader real-time audio and video conversation model rather than a dedicated translation system.

Supported inputs and outputs

The model accepts audio and image input. The supplied specifications mark native text input as unsupported, so it should not be treated as a conventional text chat model even though text appears in its responses. It also does not support native video input. Visual material can be supplied as images, while spoken content is supplied as streaming audio.

Output can be text or audio. Text may be streamed through response events such as response.text.delta or response.audio_transcript.delta, while translated speech is streamed through response.audio.delta. The result is therefore useful both for applications that display subtitles or translated transcripts and for applications that play translated speech to a listener.

CapabilitySupported detail
Audio inputYes
Image inputYes
Text inputNo, according to the supplied model specifications
Video inputNo native video input
Text outputYes, streamed in real time
Audio outputYes, translated speech
Image or video outputNo

Alibaba Cloud documentation states that the model understands 60 languages and supports speech output in 29 languages. Those are provider-documented language figures, but the exact language pairing, voice availability, regional access, and behavior should be checked against the current API documentation before deployment.

Latency and practical translation use cases

The model is designed for simultaneous interpretation, where partial results are more useful than a single response after an entire conversation has ended. Alibaba Cloud describes latency as low as approximately 2.3 seconds. That figure is a provider claim and should not be interpreted as a universal guarantee: network conditions, audio quality, message handling, language direction, and application design can all affect the observed delay.

Good use cases include multilingual video meetings, live event interpretation, translated voice calls, travel or hospitality assistance, customer-support conversations, and subtitles for live audiovisual material. For example, a meeting application could send each participant’s audio stream to the model, display translated text as it arrives, and optionally play the translated speech through a selected audio channel.

Image input gives the model a role beyond ordinary speech-to-speech translation. It can be considered for audiovisual translation workflows that include visual frames or image-based context. However, the available research does not establish a general image understanding feature set beyond its role in the translation model, so applications should validate the model against their particular visual translation task.

Context, output limits, and unsupported features

The documented context length is 53,248 tokens, and the maximum output is 4,096 tokens. These limits describe the model’s processing and response capacity; they do not mean that a live translation session should routinely accumulate a conversation of that size. For real-time applications, shorter streaming exchanges are generally easier to monitor, restart, and control.

Qwen3.8-LiveTranslate-Flash-Realtime does not support function calling, web search, structured outputs, batch inference, fine-tuning, or context caching according to the supplied model documentation. It is therefore not the right choice for an agent that must search the web, invoke business tools, return schema-validated records, or run large offline translation queues through a batch interface.

Streaming is supported, and the WebSocket Realtime API is central to the model’s design. This differs from a conventional request-and-response API: the client maintains a live connection and handles incremental events. Developers need to account for partial text, partial audio, connection state, interruption behavior, and the possibility that a real-time stream may need to be retried or re-established.

Pricing by region and modality

Pricing varies by deployment region and by the kind of tokens processed. The supplied Alibaba Cloud pricing information lists the following rates per 1 million tokens:

RegionAudio inputImage inputText outputAudio output
Singapore$7.50$0.55$20.00$30.00
China, Beijing$5.653$0.466$14.133$22.613

These are usage prices rather than a recurring subscription fee. The correct total depends on the amount and modality of input and output generated by an application. Audio output is more expensive than text output in both listed regions, so a subtitle or transcript workflow can cost less than one that synthesizes translated speech for every response. Region selection may also affect eligibility, network performance, billing, and available features.

Capability profile and trade-offs

The model’s primary strength is specialization. It is designed around low-latency translation and supports both text and spoken output, which makes it more directly useful for live interpretation than a general language model that would need additional speech recognition, translation, and speech synthesis components.

Its limitations are equally important. This is not a general-purpose reasoning or coding model. The supplied editorial assessment rates its reasoning capability at 4 out of 10 and coding capability at 2 out of 10, while rating speed at 9 out of 10 and cost efficiency at 7 out of 10. These scores are editorial evaluations, not Alibaba Cloud benchmarks or provider-published ratings. They summarize the model’s intended trade-off: strong suitability for fast translation, but limited value for complex analysis, software development, or agent workflows.

Because the model does not provide function calling, web search, structured outputs, or batch inference, it should normally sit inside a focused translation pipeline rather than serve as the central intelligence for a larger automated system. A surrounding application may still add its own business logic, storage, moderation, or routing, but those capabilities are not native features of this model.

When to choose Qwen3.8-LiveTranslate-Flash-Realtime

Choose this model when the main requirement is near-real-time translation of audio or audiovisual input and the result should arrive as text, speech, or both. It is especially appropriate when:

  • Users need live interpretation during meetings or conversations.
  • Audio must be translated continuously instead of processed as a completed recording.
  • An application needs translated speech as well as captions or transcripts.
  • Low latency is more important than deep reasoning or broad tool access.
  • The deployment can use Alibaba Cloud Model Studio’s WebSocket Realtime API.
  • Regional pricing and supported language availability fit the application’s requirements.

A different option may be more appropriate for general chat, coding, research, web-connected assistants, structured business data, or large offline translation jobs. A broader real-time multimodal model may be preferable when the application needs open-ended conversation rather than dedicated translation. The supplied research specifically identifies Qwen3.8-Omni-Flash-Realtime as a broader real-time audio and video conversation model, while Qwen3.8-LiveTranslate-Flash-Realtime remains the more focused choice when translation is the central task.

Bottom line

Qwen3.8-LiveTranslate-Flash-Realtime is a focused streaming translation model for applications that need fast multilingual audio and visual translation. Its combination of audio and image input, text and speech output, documented support for 60 languages, and approximately 2.3-second provider-reported latency makes it a practical candidate for live interpretation. The trade-off is a narrow feature set: no tools, web search, structured outputs, batch inference, fine-tuning, or native video input. It is best evaluated as a real-time translation component, not as a replacement for a general-purpose conversational or reasoning model.


Answers to Frequently Asked Questions

What are the main limitations of Qwen3.8-LiveTranslate-Flash-Realtime?
The model does not support native video input, function calling, web search, structured outputs, batch inference, fine-tuning, or context caching. It is intended as a focused real-time translation component rather than a general-purpose reasoning, coding, research, or agent model.
What is the latency and pricing of Qwen3.8-LiveTranslate-Flash-Realtime?
Alibaba Cloud reports latency as low as approximately 2.3 seconds, although actual performance depends on network conditions, audio quality, language direction, and application design. Listed pricing per 1 million tokens ranges from $5.653 to $7.50 for audio input, $0.466 to $0.55 for image input, $14.133 to $20.00 for text output, and $22.613 to $30.00 for audio output, depending on region.
What inputs and outputs does Qwen3.8-LiveTranslate-Flash-Realtime support?
The model supports streaming audio and image input, but not native text or video input. It provides streamed text output for captions or transcripts and streamed audio output for translated speech.
How many languages does Qwen3.8-LiveTranslate-Flash-Realtime support?
Alibaba Cloud states that the model understands 60 languages and supports speech output in 29 languages. Exact language pairings, voice availability, and regional access should be confirmed in the current API documentation.
What is Qwen3.8-LiveTranslate-Flash-Realtime used for?
Qwen3.8-LiveTranslate-Flash-Realtime is designed for low-latency translation of live audio and visual content. It can support multilingual meetings, live interpretation, translated voice calls, customer service, hospitality assistance, and real-time subtitles.


Sources 4
Provider

About Qwen