Qwen-Audio-TTS

qwen-audio-3.0-tts-plus

by Qwen · Current and available

Alibaba Cloud's Qwen-Audio-3.0-TTS-Plus generates expressive speech from text with multilingual and Chinese dialect support, voice cloning, and controls for emotion, tone, character, speed, volume, and style. It uses a WebSocket API and charges by input characters, making it suited to audiobooks, dubbing, content creation, customer service, and premium voice applications.

Speech Reasoning Coding
Qwen-Audio-3.0-TTS-Plus is a specialist speech-generation model in Alibaba Cloud Model Studio. Rather than generating text, images, or general-purpose answers, it converts written input into natural-sounding speech for applications such as audiobooks, dubbing, customer-service voice systems, and premium content production. Its main distinction is expressive control: developers can guide how the speech sounds, while voice cloning enables custom voice experiences.
Outputs

What qwen-audio-3.0-tts-plus can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Audio-TTS
Model type Other
Release date 2026-07-14
Status Current and available
Knowledge cutoff notes

Alibaba Cloud does not publish a model knowledge-cutoff date for this speech synthesis model. Its documented capabilities concern text-to-speech generation rather than factual question answering.

Model notes

Qwen-Audio-3.0-TTS-Plus is a dedicated speech synthesis model accessed through Alibaba Cloud Model Studio's WebSocket API. It accepts text and produces audio. The model supports instruction control for emotion, tone, character, speaking rate, volume, and synthesis style, along with Chinese dialects and multiple languages. Voice cloning is supported, while voice design is not supported. Alibaba Cloud documents more than 500 base voices generated through voice cloning. The model is intended to prioritize synthesis quality, naturalness, and expressiveness over minimum latency. No context window or maximum output-token limit is published because the service is billed by input characters rather than language-model tokens. Editorial scores are comparative estimates for a specialist TTS model and do not represent vendor benchmarks.

Cost

Model pricing

Input USD 0.20 per 10,000 input characters in Singapore/International; USD 0.19253 per 10,000 input characters in China (Beijing)
Output Not separately billed; pricing is based on input characters
Model guide

Qwen-Audio-3.0-TTS-Plus: Expressive Text-to-Speech with Voice Cloning

Qwen-Audio-3.0-TTS-Plus is Alibaba Cloud's dedicated text-to-speech model for producing expressive spoken audio. It supports multilingual and Chinese dialect synthesis, voice cloning, and controls for emotion, tone, character, speaking rate, volume, and speaking style. It is accessed through a WebSocket API and is priced by input characters rather than generated audio duration.

What is Qwen-Audio-3.0-TTS-Plus?

Qwen-Audio-3.0-TTS-Plus is Alibaba Cloud's text-to-speech model for turning written text into generated audio. It belongs to the Qwen-Audio-TTS family and is currently documented as an available model in Alibaba Cloud Model Studio. The model is designed for speech synthesis rather than general language generation: its output is audio, not a text response.

Text-to-speech, often abbreviated TTS, is the process of converting written words into spoken language. Qwen-Audio-3.0-TTS-Plus adds controls intended for expressive production, including emotion, tone, character, speaking rate, volume, and overall synthesis style. This makes it more suitable for narrated or performed speech than a basic utility voice that simply reads text aloud.

The model was released on July 14, 2026, according to the supplied Model Studio research. Availability and documentation can change, so production users should confirm the current model status and regional support in Alibaba Cloud's documentation.

Primary purpose and positioning

This model fits the specialist end of Alibaba Cloud's current AI catalog. It is not a general-purpose Qwen reasoning or coding model, and it is not intended for speech recognition, speech-to-speech conversion, image generation, or video generation. Its job is to create spoken audio from text with a relatively high level of control over delivery.

That positioning makes it useful when the sound of the voice matters as much as the words. A typical request might contain narration together with instructions such as a calm tone, an energetic delivery, a slower speaking rate, or a particular emotional character. The model can then synthesize the requested passage as audio.

Alibaba Cloud's documentation also describes voice cloning support. The provider documents more than 500 base voices generated through voice cloning. Voice cloning can help organizations create a recognizable narrator, branded assistant, or localized voice experience, but teams should obtain the necessary permissions before cloning or reproducing any person's voice.

Supported inputs and outputs

CapabilitySupported behavior
InputText
OutputGenerated audio speech
Audio inputNot documented as supported for this model
Image or video inputNot supported
Text outputNot the model's output modality
StreamingSupported through the WebSocket API
Structured text outputNot supported

The model's primary interface is a WebSocket API. WebSockets maintain a live connection between an application and the service, which is useful for receiving generated audio progressively or for building interactive voice experiences. The supplied research identifies streaming as supported, but does not provide a maximum audio duration or a maximum output-token value.

Expressive controls and voice cloning

Qwen-Audio-3.0-TTS-Plus is intended to do more than pronounce text accurately. Its documented controls cover several aspects of delivery:

  • Emotion: guide whether the delivery should sound, for example, warm, serious, excited, or calm when the requested style is supported.
  • Tone and character: shape the perceived personality or presentation of the speaker.
  • Speaking rate: request faster or slower delivery.
  • Volume: influence the loudness of the synthesized voice.
  • Speaking style: adapt the performance for narration, dialogue, dubbing, or other content formats.
  • Language and dialect: generate speech across multiple languages and Chinese dialects documented by Alibaba Cloud.

These controls are provider-documented capabilities, not a guarantee that every instruction will be interpreted identically in every passage. Voice quality and stylistic consistency can depend on the language, voice, wording, and instruction used. The research also states that voice cloning is supported while voice design is not. In other words, the documented feature is based on reproducing or adapting an existing voice rather than creating an entirely new voice from a textual description.

Languages, dialects, and practical use cases

Multilingual and dialect support is one of the model's main practical advantages. It can be considered for content that needs the same production workflow across different languages or regional varieties, although users should test pronunciation, names, accents, and emotional delivery in each target language before publishing.

Supported use cases include:

  • Audiobooks: produce narrated chapters with controlled pacing and characterful delivery.
  • Film and video dubbing: create localized speech and adjust tone or timing for different scenes.
  • Content creation: generate narration for educational videos, podcasts, explainers, and social media content.
  • Customer service: build spoken prompts, notifications, or voice-agent responses where a consistent voice is important.
  • Premium voice applications: offer branded voices or customized voice experiences through an application.

For high-stakes customer communications, organizations should review generated audio for pronunciation, unintended emotional cues, and compliance with disclosure or consent requirements. The model's ability to synthesize a voice does not remove the need for human review.

Pricing and API access

Alibaba Cloud prices Qwen-Audio-3.0-TTS-Plus according to input characters rather than separately charging for generated audio output. The supplied pricing information is:

RegionPrice
Singapore and internationalUSD 0.20 per 10,000 input characters
China, Beijing regionUSD 0.19253 per 10,000 input characters

Generated audio output is not separately billed in the supplied pricing description. The regional prices should not be treated as interchangeable: the applicable amount depends on the Model Studio region and account configuration. Before deployment, verify the current price, eligible region, quotas, and any account requirements in Alibaba Cloud's pricing documentation.

Character-based billing can make costs easier to estimate for a known script. For example, a production team can approximate the synthesis charge from the number of characters in an audiobook chapter or batch of dialogue. The final application cost may still include storage, delivery, processing, voice review, and other cloud-service charges that are outside the model price.

Limits and unsupported capabilities

Alibaba Cloud does not publish a context-window size or maximum output-token limit for this model in the supplied research. That is expected for a speech service billed by input characters rather than language-model tokens, but it means developers should not assume that arbitrarily long scripts can be submitted in one request. Splitting long content into chapters, scenes, or manageable segments is a practical integration pattern, subject to the service's current request and quota limits.

The model is not a suitable choice when the application needs:

  • general-purpose text generation or factual question answering;
  • reasoning, planning, or coding from the model itself;
  • speech recognition or transcription from an audio recording;
  • speech-to-speech transformation;
  • image or video generation;
  • native JSON or other structured text output; or
  • tool calling and function execution.

The research lists reasoning and coding scores of 1, but these are editorial comparative scores for a specialist TTS model, not Alibaba Cloud benchmark results. They should be read as an indication that reasoning and coding are outside the model's intended role, not as measurements of speech quality.

Speed, quality, and cost trade-offs

The supplied notes characterize Qwen-Audio-3.0-TTS-Plus as prioritizing synthesis quality, naturalness, and expressiveness over minimum latency. That is an important trade-off for interactive systems. A live voice interface that must respond almost immediately may prefer a lower-latency speech model, while an audiobook, dubbed video, or premium narration workflow may benefit more from expressive delivery and voice consistency.

Streaming through WebSockets can improve the user experience by allowing an application to handle audio as it becomes available, but streaming does not necessarily mean that the entire synthesis process is instantaneous. Teams should test time to first audio, completion time, pronunciation accuracy, and consistency across representative scripts.

The editorial cost score of 8 and speed score of 8 in the supplied research are comparative estimates, not provider-published ratings. The documented price per input character is the verifiable cost information; actual value depends on how much editing, retaking, localization, and human review a project requires.

When to choose Qwen-Audio-3.0-TTS-Plus

Choose this model when the central requirement is expressive text-to-speech with multilingual or dialect coverage, voice cloning, and controllable delivery. It is particularly well matched to narrated media, dubbing, branded voice experiences, and applications where a plain readout would sound insufficiently natural or distinctive.

Another speech model may be more appropriate when the top priority is the lowest possible latency, when only basic text-to-speech is needed at large scale, or when the application requires a different voice-creation workflow. A speech-recognition model is the better option for converting audio into text, while a general-purpose language model is needed for reasoning, coding, or generating the script itself. Those models can be combined with Qwen-Audio-3.0-TTS-Plus in a larger pipeline, but they are separate capabilities and should not be attributed to this TTS model.

Overall assessment

Qwen-Audio-3.0-TTS-Plus is best understood as a production-oriented voice synthesis component rather than an all-purpose AI model. Its strongest documented characteristics are expressive controls, multilingual and Chinese dialect support, voice cloning, and WebSocket streaming. Its pricing is based on input characters, with no separately billed generated audio output in the supplied description.

The main evaluation questions are therefore practical: does the chosen voice pronounce the project's vocabulary correctly, does the emotional direction remain consistent, is the regional price acceptable, and is the model's latency suitable for the application? For teams that answer yes, it offers a focused route to customized spoken audio. For tasks involving text reasoning, transcription, tool use, or other modalities, a separate model should handle those parts of the workflow.


Answers to Frequently Asked Questions

Does Qwen-Audio-3.0-TTS-Plus support streaming and audio input?
The model supports audio streaming through Alibaba Cloud's WebSocket API, which can help applications receive generated speech progressively. Audio input is not documented as supported, so the model should not be treated as a speech-recognition or speech-to-speech system.
How much does Qwen-Audio-3.0-TTS-Plus cost?
Pricing is based on input characters. The supplied pricing lists USD 0.20 per 10,000 input characters in Singapore and international regions, and USD 0.19253 per 10,000 input characters in the China Beijing region. Generated audio is not separately billed in the supplied pricing description.
Does Qwen-Audio-3.0-TTS-Plus support voice cloning?
Yes. Alibaba Cloud documents voice cloning support and more than 500 base voices generated through voice cloning. Users should obtain appropriate permission before reproducing or adapting any person's voice.
What expressive controls does Qwen-Audio-3.0-TTS-Plus offer?
The model supports controls for emotion, tone, character, speaking rate, volume, speaking style, language, and dialect. It can be directed toward delivery styles such as calm, energetic, serious, narrated, or conversational, although results may vary by voice, language, and prompt.
What is Qwen-Audio-3.0-TTS-Plus used for?
Qwen-Audio-3.0-TTS-Plus is an Alibaba Cloud text-to-speech model that converts written text into generated audio. It is designed for expressive narration, dubbing, audiobooks, content creation, customer-service prompts, and branded voice applications.


Sources 6
Provider

About Qwen