MiMo-V2.5-TTS

MiMo-V2.5-TTS-VoiceClone

by Xiaomi HyperAI · Current; limited-time free access

Xiaomi MiMo-V2.5-TTS-VoiceClone is a hosted speech model that clones a speaker from a short MP3 or WAV sample and generates expressive speech from text. It supports Chinese and English reference audio, style and emotion controls, and OpenAI-compatible API access. The model is currently free for a limited time, but its streaming interface does not provide true low-latency audio, and it is not intended for reasoning, coding, tool use, or local deployment.

Speech
MiMo-V2.5-TTS-VoiceClone is Xiaomi's specialized zero-shot voice-cloning model. Developers provide a short reference recording and text, and the model generates speech that imitates the reference speaker without requiring voice-specific training or fine-tuning. It is aimed at custom narration, character dialogue, accessibility voices, interactive applications, and other audio workflows where a consistent speaker identity matters.
Outputs

What MiMo-V2.5-TTS-VoiceClone can produce

Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiMo-V2.5-TTS
Model type Other
Context window 8K tokens
Maximum output 8K tokens
Release date 2026-05-28
Status Current; limited-time free access
Knowledge cutoff notes

Xiaomi does not publish a separate knowledge cutoff for this speech synthesis model. Its behavior is controlled through supplied text, reference audio, and style instructions rather than a conventional factual knowledge cutoff.

Model notes

Canonical API model ID: mimo-v2.5-tts-voiceclone. Accepts text plus a Base64-encoded MP3 or WAV reference sample; Xiaomi documents a 10 MB limit for the encoded sample. The cloned voice supports natural-language style instructions and audio-tag controls, including emotional and delivery changes. Chinese and English reference audio are supported. The model does not support built-in voices, singing mode, or text-based voice design. Xiaomi's documentation lists streaming output, but the current voice-cloning guide says low-latency streaming is unavailable and streaming requests operate in compatibility mode, returning the completed result after inference. The model is hosted through Xiaomi's OpenAI-compatible API and is not documented as open-weight or downloadable. TTS models are currently free for a limited time and do not consume Token Plan credits during the promotional period.

Cost

Model pricing

Input Free for a limited time
Output Free for a limited time
Model guide

MiMo-V2.5-TTS-VoiceClone: Zero-Shot Custom Voice Synthesis

Xiaomi MiMo-V2.5-TTS-VoiceClone is a hosted text-to-speech model that creates expressive speech in a cloned speaker's voice from a short MP3 or WAV sample. It supports Chinese and English reference audio, natural-language style instructions, emotion and delivery controls, and Xiaomi's OpenAI-compatible API, but it is not designed for true low-latency streaming, singing, local deployment, or text-based reasoning.

What is MiMo-V2.5-TTS-VoiceClone?

MiMo-V2.5-TTS-VoiceClone is a hosted speech-synthesis model from Xiaomi MiMo. Its defining capability is zero-shot voice cloning: a developer supplies a short recording of a speaker, and the model uses that sample to generate new spoken audio from text. “Zero-shot” means that the application does not need to train or fine-tune a separate model for each voice.

The model is listed in Xiaomi's MiMo V2.5 TTS family under the API identifier mimo-v2.5-tts-voiceclone. It is accessed as an online service through Xiaomi's OpenAI-compatible API rather than as a downloadable or self-hosted model checkpoint.

Voice cloning here involves more than copying a speaker's general timbre. According to Xiaomi's documentation, the generated voice can also follow instructions about emotion, pace, tone, delivery, and performance style. This makes the model suitable for producing different performances while retaining a recognizable target voice.

How the voice-cloning workflow works

A typical request contains three important elements: a reference recording, the text to speak, and optional instructions describing how the text should be delivered. The reference recording is supplied as a Base64-encoded data URL in the audio voice parameter. Xiaomi documents MP3 and WAV as supported reference formats, with a 10 MB limit for the encoded audio payload.

The target text is passed to the speech-synthesis request, while style or performance directions can be included separately. For example, an application could request that a cloned voice read a sentence slowly and calmly, deliver dialogue with excitement, or perform a line as a fictional character. The generated response contains Base64-encoded speech audio.

The provider's examples use the OpenAI Python client together with Xiaomi's API base URL. This compatibility can reduce integration work for developers already familiar with OpenAI-style client libraries, but it does not turn the model into a general-purpose conversational or reasoning system. Its job is speech generation.

Supported inputs, outputs, and limits

SpecificationVerified detail
Model typeText-to-speech and voice cloning
Text inputSupported
Reference audio inputMP3 or WAV, supplied as Base64 data
Reference payload limit10 MB
Audio outputGenerated speech audio
Reference languagesChinese and English are documented
Context length8,000 tokens in Xiaomi's model specification
Maximum output8,000 tokens in Xiaomi's model specification
Image and video inputNot supported
Text outputNot the model's native output

Xiaomi lists an 8K-token context length and an 8K-token maximum output in its general model specifications. These values should be interpreted cautiously for a speech model: they are token-based platform limits, not a direct promise of a particular number of minutes of audio. Actual usable duration can depend on the text, language, request format, and service behavior.

The model accepts multimodal input in the limited sense that it processes both text and reference audio. Its output is audio only. It does not generate images or video, return embeddings, execute code, or provide a separate text-based reasoning answer.

Expression and performance control

The practical advantage of MiMo-V2.5-TTS-VoiceClone is the combination of speaker identity and controllable delivery. The reference sample establishes whose voice the system should imitate, while instructions can influence how the generated line sounds.

  • Emotion: Request a calmer, happier, more urgent, or otherwise emotionally directed performance where the model follows the instruction.
  • Pacing: Ask for faster or slower delivery to suit narration, dialogue, or instructional content.
  • Role and character direction: Describe a performance style or character presentation while retaining the cloned voice.
  • Audio tags: Xiaomi documents inline audio-tag controls that provide additional performance guidance.
  • Language coverage: Chinese and English reference audio support is documented by Xiaomi.

These controls make the model more flexible than a system that simply reads text in a fixed cloned voice. They do not guarantee perfect compliance with every instruction, and the supplied research does not provide benchmark results or a quantified similarity score. Claims about high fidelity should therefore be treated as provider positioning rather than as an independently verified performance measurement.

Pricing and availability

Xiaomi currently lists the TTS series as free for a limited time. The model is also included in the MiMo V2.5 TTS family supported by Xiaomi's Token Plan, where Xiaomi states that TTS models do not consume package credits during the promotional period.

This is promotional availability rather than a permanent price guarantee. The free period, eligibility rules, quotas, or billing treatment may change, so production users should check Xiaomi's current model and pricing documentation before committing to a large deployment. No stable per-character, per-minute, or per-request price is established in the supplied research.

Access requires use of the hosted Xiaomi MiMo platform and its API credentials. The model is not documented as an open-weight release, so organizations that need offline processing, private infrastructure, or full control over model weights should consider a different type of solution.

Streaming and latency limitations

The API accepts streaming requests, but this capability needs an important qualification. Xiaomi's current voice-cloning documentation says that true low-latency incremental streaming is not available for this model. Streaming operates in compatibility mode and may return the completed result only after inference has finished.

That distinction matters for interactive products. A request may use a streaming-shaped interface without delivering audio progressively as it is generated. The model can still be useful for prerecorded responses, generated clips, narration, and applications that can wait for a complete result, but it is a weaker fit for natural, interruption-heavy voice conversations where the first audio must arrive quickly.

Reasoning, coding, and tool support

MiMo-V2.5-TTS-VoiceClone is a speech-generation model, not a general-purpose language model. It has no documented reasoning capability score, and it should not be selected for factual analysis, planning, mathematics, or long-form text generation as a primary task.

Coding support is not documented and is not relevant to the model's intended role. The model does not provide code execution, function calling, web search, or other tool-use capabilities in the supplied specifications. An application that needs dialogue generation, retrieval, moderation, or business logic should perform those tasks with separate software or another model, then send the final text to VoiceClone for synthesis.

Likewise, JSON or structured-output generation is not a native purpose of this model. Although the surrounding API may use structured request and response fields, that should not be confused with a general JSON-mode capability.

Main strengths and trade-offs

  • Custom voice without per-voice training: A short reference sample is intended to be enough to begin cloning, reducing the setup burden compared with training a dedicated voice model.
  • Expressive delivery: Style instructions, emotion controls, and audio tags provide more control than fixed-voice text-to-speech.
  • Developer access: The OpenAI-compatible API and hosted delivery provide a practical integration path for applications that already use compatible client patterns.
  • Promotional cost: Xiaomi currently lists the TTS series as free for a limited time, with no Token Plan credit consumption during the stated promotion.
  • Limited real-time suitability: Compatibility-mode streaming does not provide the true low-latency behavior required by many live voice agents.
  • Hosted-service dependency: The model is not documented as downloadable or open-weight, so the application depends on Xiaomi's service availability, policies, and pricing.
  • Narrow task focus: It generates speech but does not replace a text model, reasoning model, tool system, or application orchestration layer.

Best use cases

MiMo-V2.5-TTS-VoiceClone is a strong fit when the application needs a recognizable custom speaker and can tolerate hosted, non-instant generation. Suitable examples include:

  • Localized narration using a consistent voice identity.
  • Character dialogue for games, prototypes, and media production.
  • Personalized accessibility or reading voices, with the speaker's permission.
  • Voice prototypes for presentations, educational material, or product testing.
  • Generated audio responses in applications where a short wait for completed synthesis is acceptable.
  • Multilingual or bilingual experiments involving documented Chinese and English reference audio support.

Organizations should obtain appropriate permission before cloning a person's voice and should disclose synthetic or cloned speech where required. Consent, identity protection, and abuse prevention are especially important when a recognizable real person's recording is used.

When to choose this model

Choose MiMo-V2.5-TTS-VoiceClone when the priority is zero-shot custom-voice synthesis, expressive narration, or character performance rather than general intelligence. It is particularly attractive when a short reference recording is available, the application can send audio to a hosted API, and the current promotional pricing is useful.

Another option may be more appropriate in several situations. Choose a real-time speech system when the product requires genuinely incremental, low-latency audio for live conversation. Choose a fixed-voice TTS service when custom identity is unnecessary and predictable voice selection is more important. Choose a locally deployable or open-weight model when sensitive recordings cannot leave private infrastructure. Choose a general language model alongside a TTS system when the application must reason, retrieve information, call tools, or generate complex dialogue before speaking.

Within Xiaomi's lineup, MiMo-V2.5-TTS-VoiceClone should be understood as the custom-reference voice option in the MiMo V2.5 TTS family. Other TTS models may serve different purposes, such as built-in voice selection, text-based voice design, or singing, but those functions are not documented for this VoiceClone model.

Bottom line

MiMo-V2.5-TTS-VoiceClone provides a focused way to turn text into expressive speech that imitates a supplied speaker. Its most important differentiators are zero-shot voice cloning, style and emotion control, Chinese and English reference support, and straightforward hosted API access. Its main constraints are the 10 MB reference payload limit, the platform's 8K-token specifications, promotional rather than guaranteed pricing, lack of documented reasoning or tool use, and the absence of true low-latency streaming. For custom narrated or character audio, it is a practical candidate; for live voice agents, private local inference, or general-purpose AI work, another architecture is likely to be a better fit.


Answers to Frequently Asked Questions

When should developers choose MiMo-V2.5-TTS-VoiceClone?
Developers should choose it when they need expressive custom-voice synthesis for narration, character dialogue, accessibility, or prototypes and can use a hosted API. It is less suitable for private offline deployment, genuinely real-time voice interaction, or applications that also require reasoning, retrieval, coding, tool use, or dialogue generation.
Does MiMo-V2.5-TTS-VoiceClone support real-time streaming?
The API accepts streaming-shaped requests, but Xiaomi's documentation indicates that true low-latency incremental streaming is not available. Responses may be returned only after the complete inference finishes, making the model better suited to prerecorded audio and narration than live conversational agents.
What are the input limits and supported languages for MiMo-V2.5-TTS-VoiceClone?
The model supports MP3 and WAV reference audio with a 10 MB limit for the encoded reference payload. Xiaomi documents Chinese and English reference audio support, an 8,000-token context length, and an 8,000-token maximum output.
What is MiMo-V2.5-TTS-VoiceClone?
MiMo-V2.5-TTS-VoiceClone is Xiaomi MiMo's hosted text-to-speech model for zero-shot voice cloning. It uses a short speaker recording to generate new speech from text without requiring per-voice training or fine-tuning.
How does voice cloning with MiMo-V2.5-TTS-VoiceClone work?
An application submits a reference MP3 or WAV recording as a Base64-encoded data URL, along with the text to synthesize and optional instructions for emotion, pacing, tone, or performance style. The API returns generated speech audio, also encoded as Base64.


Sources 5
Provider

About Xiaomi HyperAI