MiMo-V2.5-TTS

MiMo-V2.5-TTS-VoiceDesign

by Xiaomi HyperAI · Current; limited-time free access

Xiaomi MiMo-V2.5-TTS-VoiceDesign generates speech from voices described in natural language. It supports control over vocal characteristics, accents, emotions, rhythm, and performance style without requiring a preset voice or reference recording, but it does not provide voice cloning, singing mode, or confirmed true low-latency streaming.

Speech
MiMo-V2.5-TTS-VoiceDesign takes a written description of a desired voice and uses it to synthesize spoken text in that style. Unlike a fixed voice catalog, it is designed to create a new voice for each use case without cloning a real speaker or uploading a reference sample. Xiaomi provides the model through its OpenAI-compatible MiMo API.
Outputs

What MiMo-V2.5-TTS-VoiceDesign can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

6/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiMo-V2.5-TTS
Model type Other
Context window 8K tokens
Maximum output 8K tokens
Release date 2026-04-23
Status Current; limited-time free access
Model notes

Canonical API model ID is mimo-v2.5-tts-voicedesign. The model creates a new voice from a text description and does not require a reference audio sample. Xiaomi states that built-in voices, voice cloning, and singing mode are not supported for this model. The API uses an OpenAI-compatible Chat Completions interface: the user message contains the voice description and the assistant message contains the target speech text. The optional optimize_text_preview parameter can allow the assistant message to be omitted. Streaming is supported through a compatibility interface, but Xiaomi's current documentation says true low-latency streaming is not yet available for this model. The model is currently listed with 100 RPM and 10M TPM limits. Pricing is promotional and described as free for a limited time.

Cost

Model pricing

Input Free for a limited time
Output Free for a limited time
Model guide

MiMo-V2.5-TTS-VoiceDesign: Create Custom Voices from Text Descriptions

MiMo-V2.5-TTS-VoiceDesign is Xiaomi's speech-synthesis model for generating custom synthetic voices from natural-language descriptions. It lets developers specify characteristics such as age, accent, vocal texture, emotion, speed, rhythm, and performance style without supplying a reference audio recording. The model outputs speech audio through Xiaomi's MiMo API and is aimed at narration, characters, podcasts, ASMR, games, assistants, and other creative audio applications.

What is MiMo-V2.5-TTS-VoiceDesign?

MiMo-V2.5-TTS-VoiceDesign is Xiaomi's text-to-speech model for generating speech with a voice defined in natural language. Its distinguishing feature is voice design: instead of selecting a preset voice or providing a recording to imitate, the developer describes the intended speaker and performance.

A description might request a calm elderly storyteller, a gravelly middle-aged narrator with a regional accent, or a soft, slow ASMR performer. The model then synthesizes the requested speech using those characteristics. The canonical API model identifier is mimo-v2.5-tts-voicedesign.

The model belongs to Xiaomi's MiMo-V2.5-TTS series. Within that series, VoiceDesign serves applications that need a newly specified voice. It is distinct from the standard MiMo-V2.5-TTS model, which uses built-in voices, and from MiMo-V2.5-TTS-VoiceClone, which is intended for voice cloning from an audio reference.

How the voice-design workflow works

The API uses an OpenAI-compatible Chat Completions interface with a model-specific message arrangement. The user message contains the voice description, while the assistant message contains the text that should be spoken. In practical terms, one message explains who should sound like what, and the other provides what that speaker should say.

A useful voice description should focus on audible characteristics and performance direction. Xiaomi's guidance supports attributes such as:

  • Age and gender
  • Accent or regional speech characteristics
  • Vocal texture, including softness, roughness, breathiness, or clarity
  • Emotion and mood
  • Speaking speed and rhythm
  • Role, setting, and performance style

Descriptions should generally use one to four clear sentences and avoid contradictory instructions. For example, asking for a slow, meditative bedtime narrator is more coherent than combining mutually conflicting directions such as “very fast,” “long pauses,” and “extremely relaxed.” Xiaomi also advises describing the voice itself rather than post-processing effects such as echo, reverb, equalization, or compression.

Example of a practical voice description

A suitable instruction could be: “A warm, elderly woman from northern England, speaking slowly and clearly with a gentle smile in her voice, like a grandmother telling a child a comforting bedtime story.” The accompanying speech text should fit that direction. A quiet bedtime monologue is a better match than an energetic sports announcement.

Chinese and English voice descriptions are supported according to Xiaomi's documentation. The optional optimize_text_preview parameter can ask the service to polish the target text; when it is enabled, Xiaomi documents that the assistant message may be omitted.

Capabilities and audio output

VoiceDesign is designed for natural-language control of synthetic speech rather than general-purpose text generation. It can vary the perceived speaker and delivery across supported combinations of age, gender, accent, vocal texture, emotional tone, speed, rhythm, and context.

The model produces audio output. Xiaomi's examples document common audio formats such as WAV and PCM16 output for streaming-compatible calls. The documented PCM16 example uses 24 kHz, mono PCM16LE audio, which can be useful for applications that need a conventional raw audio representation.

The model's recorded context limit is 8,192 tokens, with a maximum output limit of 8,192 tokens. These limits are relevant to the API request and generated response, although actual practical speech length can also depend on the selected audio format, request structure, and service behavior.

SpecificationDocumented value
ProviderXiaomi
Model IDmimo-v2.5-tts-voicedesign
Primary inputText voice description and synthesis text
Primary outputGenerated speech audio
Context length8,192 tokens
Maximum output8,192 tokens
Image, video, or audio inputNot documented for this model; voice design does not require reference audio
Text outputNot the model's primary output; it is a speech-synthesis model

API access and streaming behavior

MiMo-V2.5-TTS-VoiceDesign is available through Xiaomi's MiMo API rather than as a standalone consumer voice-design application. The API is described as OpenAI-compatible, but applications still need to follow Xiaomi's model-specific message format and audio parameters.

The model exposes a streaming-compatible interface, but this should not be confused with true low-latency streaming. Xiaomi's current speech-synthesis documentation says that VoiceDesign streaming requests can operate in compatibility mode and may return the result only after inference has finished. An application that needs audio chunks immediately while speech is being generated should therefore treat this model as unsuitable for that requirement.

This distinction matters for interactive assistants, live dialogue, and phone-style conversations. The interface may accept a streaming setting, yet the resulting experience may still be effectively request-and-wait rather than incremental playback.

Main strengths and limitations

Where the model is strong

  • Text-based customization: Developers can define a voice without recording a speaker or searching through a fixed catalog.
  • Broad expressive control: Voice descriptions can include accent, age, emotion, vocal texture, speed, rhythm, and role.
  • Useful for fictional voices: It is well suited to characters, narration, games, audio prototypes, and other situations where an original voice is preferable to an exact replica.
  • Accessible API workflow: The OpenAI-compatible interface can be easier to integrate for applications already built around Chat Completions-style requests.
  • Promotional pricing: Xiaomi currently describes the TTS series as free for a limited time, although that offer should not be treated as a permanent price.

Important limitations

  • No reference-audio voice cloning: VoiceDesign does not clone a real speaker from an audio sample. Xiaomi identifies VoiceClone as the relevant sibling model for that use case.
  • No built-in voice selection: It is not intended to provide a stable catalog of predefined voices. The standard MiMo-V2.5-TTS model is the more appropriate comparison when preset voices are needed.
  • No documented singing mode: The singing capability documented for the standard MiMo-V2.5-TTS model is not supported here.
  • Not true low-latency streaming: Compatibility-mode streaming may wait for the completed inference result.
  • Prompt consistency matters: Contradictory or vague descriptions can make the requested voice less predictable.
  • Limited modality scope: This is an audio-generation model, not a model for image, video, coding, or general tool execution.

Reasoning, coding, and tool support

MiMo-V2.5-TTS-VoiceDesign is not presented as a reasoning or coding model. Its purpose is to convert text and voice-direction instructions into speech audio. Xiaomi's supplied model information does not document a reasoning capability, code-generation capability, function calling, web search, or other tool-use feature for this model.

That does not prevent an application from using a separate language model to prepare scripts or voice descriptions before sending them to VoiceDesign. It does mean that those planning, coding, and tool features should be treated as external application components rather than capabilities of the speech model itself.

Pricing, speed, and cost considerations

Xiaomi currently lists the model as free for a limited time. No permanent paid per-token price is established in the supplied documentation, so developers should verify the current MiMo pricing page before designing a long-term cost model.

The documented service limits are 100 requests per minute and 10 million tokens per minute. These are rate limits, not a guarantee of generation speed. In editorial terms, the model's main cost advantage is its current promotional availability, while its main operational trade-off is that compatibility-mode streaming may prevent the low-latency experience expected from a live voice system.

For batch narration, character audio, and creative testing, waiting for a complete result may be acceptable. For live conversation, users may prefer a speech system with confirmed incremental audio delivery even if it offers less flexible voice design or has a different price structure.

Best use cases

MiMo-V2.5-TTS-VoiceDesign is a strong fit when the application needs an original voice that can be described in words rather than selected from a fixed list. Suitable uses include:

  • Game and interactive-fiction characters
  • Audiobook, podcast, and documentary-style narration
  • Short-form video voiceovers
  • Branded or fictional assistants
  • ASMR and atmospheric audio experiments
  • Educational and accessibility content
  • Role-play and storytelling experiences
  • Rapid creative prototypes where recording a voice actor is impractical

It is less suitable for exact replication of a real person, a predictable preset-voice catalog, singing, or applications that require audio to arrive in low-latency incremental chunks.

When to choose this model

Choose MiMo-V2.5-TTS-VoiceDesign when the central requirement is custom voice creation from a written description. It is especially attractive when a project needs multiple fictional voices, unusual accents or textures, or rapid experimentation without collecting reference recordings.

Choose the standard MiMo-V2.5-TTS model instead when built-in voices or its documented singing mode are more important. Consider MiMo-V2.5-TTS-VoiceClone when the goal is to reproduce a voice from an audio sample. For a live conversational product, evaluate whether the model's compatibility-mode streaming behavior meets the required response experience before committing to it.

Overall assessment

MiMo-V2.5-TTS-VoiceDesign occupies a specific position in Xiaomi's speech lineup: it prioritizes natural-language control over a newly designed voice rather than voice cloning, preset selection, or singing. Its strongest practical advantage is the ability to move from a written creative brief to generated speech without a reference recording. Its most important limitation is that the API's streaming compatibility does not currently provide the same low-latency behavior as true incremental audio generation.

For narration, fictional characters, podcasts, ASMR, and exploratory audio production, that trade-off can be worthwhile, particularly while Xiaomi's limited-time free availability remains in effect. For real-time dialogue, exact speaker identity, or preset-voice consistency, another model type may be a better fit.


Answers to Frequently Asked Questions

What are the best use cases for MiMo-V2.5-TTS-VoiceDesign?
The model is well suited to fictional characters, games, interactive fiction, audiobooks, podcasts, documentaries, video voiceovers, branded assistants, ASMR, educational content, accessibility applications, role-play, and rapid audio prototyping. It is less suitable for exact voice replication, preset voice catalogs, singing, or applications requiring immediate audio chunks.
Does MiMo-V2.5-TTS-VoiceDesign support real-time audio streaming?
It provides a streaming-compatible API interface, but Xiaomi documents that compatibility-mode requests may return audio only after inference is complete. It should therefore not be assumed to provide true low-latency, incremental audio suitable for live conversations.
Does MiMo-V2.5-TTS-VoiceDesign clone voices from audio recordings?
No. MiMo-V2.5-TTS-VoiceDesign creates a new voice from a written description and does not require reference audio. For cloning a real speaker from an audio sample, Xiaomi identifies MiMo-V2.5-TTS-VoiceClone as the appropriate model.
What is MiMo-V2.5-TTS-VoiceDesign?
MiMo-V2.5-TTS-VoiceDesign is Xiaomi's text-to-speech model for creating synthetic speech from a natural-language voice description. Its API model identifier is "mimo-v2.5-tts-voicedesign".
How do you create a custom voice with MiMo-V2.5-TTS-VoiceDesign?
Provide a voice description in the user message and the text to be spoken in the assistant message through Xiaomi's OpenAI-compatible Chat Completions API. Describe characteristics such as age, gender, accent, vocal texture, emotion, speed, rhythm, and performance style in one to four clear, non-contradictory sentences.


Sources 6
Provider

About Xiaomi HyperAI