MiniMax Speech

Speech-02-HD

by MiniMax · Legacy or older generation; still accessible through some partner platforms, but not listed as a core model on MiniMax's current global pricing page

High-fidelity MiniMax speech model for expressive multilingual text-to-speech, zero-shot voice cloning, narration, audiobooks, voiceovers, and digital characters. It supports multiple audio formats and streaming, but its current direct pricing and long-term availability are less certain because newer MiniMax speech generations are now emphasized.

Speech Reasoning Coding
MiniMax Speech-02-HD is the high-fidelity member of the Speech-02 text-to-speech family. It is designed for natural-sounding, expressive speech rather than general-purpose reasoning or conversational text generation. The model supports multilingual narration, zero-shot voice cloning, emotional delivery, and production-oriented audio formats, making it a candidate for audiobooks, voiceovers, education, advertising, and character voices. However, it is an older model in MiniMax's rapidly changing speech lineup, and current access may depend on a partner platform or regional availability.
Outputs

What Speech-02-HD can produce

Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
6/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family MiniMax Speech
Model type Other
Context window 10K tokens
Release date 2025-05-16
Status Legacy or older generation; still accessible through some partner platforms, but not listed as a core model on MiniMax's current global pricing page
Knowledge cutoff notes

Speech synthesis models do not use a conventional public knowledge cutoff in the same way as text language models. No exact provider-published knowledge cutoff was found.

Model notes

The canonical model is commonly identified as speech-02-hd or Speech-02-HD. MiniMax introduced Speech-02 on May 16, 2025 and described Speech-02-HD as the high-fidelity variant for voiceovers and audiobooks. The official technical report identifies the evaluated MiniMax-Speech system with Speech-02-HD and documents a learnable speaker encoder, zero-shot voice cloning, cross-lingual synthesis, and support for 32 languages. The 10,000 value is a maximum input length reported by Alibaba Cloud Model Studio and should be interpreted as an input-character limit rather than a token context window. MiniMax's current core pricing page emphasizes speech-2.8-hd, so direct first-party pricing and lifecycle status for Speech-02-HD are not fully verified.

Cost

Model pricing

Input CNY 3.5 per 10,000 characters on Alibaba Cloud Model Studio in China; current direct MiniMax pricing for this exact model is not verified
Model guide

MiniMax Speech-02-HD: High-Fidelity Voice Cloning and Multilingual Text-to-Speech

MiniMax Speech-02-HD is a quality-focused text-to-speech model for expressive multilingual narration, voiceovers, audiobooks, digital characters, and zero-shot voice cloning. It supports text input, optional reference audio for voice identity, and audio output in formats including WAV, MP3, FLAC, and PCM. Its main trade-off is that it prioritizes fidelity and expressiveness over the lowest latency, while its current availability and direct MiniMax pricing are less certain because newer Speech generations are now emphasized in the provider's catalog.

What is MiniMax Speech-02-HD?

MiniMax Speech-02-HD is a text-to-speech model from MiniMax that converts written text into spoken audio. The HD designation identifies it as the quality-oriented variant of the Speech-02 family. It is intended for cases where pronunciation, vocal identity, rhythm, and expressive delivery matter more than achieving the shortest possible generation time.

The model was introduced with the Speech-02 family on May 16, 2025. MiniMax presented Speech-02-HD for voiceovers, audiobooks, narration, digital characters, education, advertising, and other audio-production tasks. The related Speech-02-Turbo variant was positioned for faster interactive generation, so the basic distinction is straightforward: Speech-02-HD favors fidelity and expressive quality, while Turbo favors speed.

Speech-02-HD should not be confused with a general-purpose language model. It does not provide text reasoning, coding assistance, web search, function calling, image generation, video generation, embeddings, or structured JSON output. Its purpose is speech synthesis and voice adaptation.

Core capabilities and supported inputs

Text-to-speech generation

The primary workflow is text input followed by generated audio. A user can supply narration, dialogue, educational material, advertising copy, or other written content and receive a spoken rendition. The model is designed to preserve natural rhythm and contextual intonation rather than simply reading every sentence in a flat voice.

Zero-shot voice cloning

Speech-02-HD supports zero-shot voice cloning. In practical terms, a short reference recording can provide the vocal identity for a new utterance without requiring a transcript of the reference audio or a conventional speaker-specific training process. This makes it possible to prototype a personalized narrator or character voice relatively quickly.

MiniMax's technical report describes a learnable speaker encoder. The encoder extracts information about the reference speaker and uses it to guide new speech. The research also describes cross-lingual voice cloning, which is useful when the target text is in a different language from the reference recording. Voice cloning should still be used only with appropriate permission from the speaker.

Expressive and controllable delivery

The model is intended to produce emotional variation, natural timing, and contextual intonation. MiniMax describes automatic and manual emotional control, while documented API controls include voice selection and adjustments for pitch, speed, and volume. These controls can help distinguish a calm instructional reading from a dramatic character performance or a more energetic commercial voiceover.

Multilingual speech

MiniMax reports support for 32 languages and a broad range of accents. The documented evaluation and language list includes Chinese, English, Cantonese, Japanese, Korean, Arabic, Spanish, French, Vietnamese, Thai, Turkish, Indonesian, Portuguese, German, Russian, Ukrainian, Polish, Romanian, Greek, Czech, Finnish, and Hindi, among others. Exact pronunciation quality can vary by language, accent, text, and delivery style, so the language count should be treated as a provider-reported capability rather than a guarantee of identical quality in every language.

Audio output, formats, and streaming

Speech-02-HD produces audio rather than text. The Speech-02 family documentation lists FLAC, WAV, MP3, and PCM output. These formats cover common production needs: WAV and PCM are useful when preserving uncompressed audio is important, while MP3 is more convenient for distribution and smaller files. FLAC provides lossless compression.

The family also supports real-time streaming. Streaming can make generated speech available progressively instead of requiring the complete audio file before playback. Nevertheless, Speech-02-HD is primarily positioned as a quality-focused model, not as MiniMax's lowest-latency option for highly interactive conversation. For an application where immediate turn-taking is more important than maximum fidelity, a faster speech model may be more appropriate.

Technical specifications at a glance

SpecificationSpeech-02-HD
ProviderMiniMax
Model familyMiniMax Speech
Release dateMay 16, 2025
Primary functionText-to-speech and zero-shot voice cloning
InputText, with reference audio used for voice cloning workflows
OutputAudio speech
Reported languages32
Audio formatsFLAC, WAV, MP3, and PCM
Maximum input length10,000 characters, as reported by Alibaba Cloud Model Studio
Maximum output tokensNot reported
StreamingSupported by the Speech-02 family
Tool or function callingNot supported
Structured JSON outputNot supported
Fine-tuningNot reported as supported

The 10,000 figure is an input-character limit reported for Alibaba Cloud Model Studio. It should not be interpreted as a token context window, because speech synthesis documentation describes the limit in characters. A provider-published maximum output length was not found. Output duration will depend on the amount of text, language, speaking rate, and other generation settings.

Quality, speed, and cost trade-offs

The main trade-off is between high-fidelity delivery and low latency. Speech-02-HD is the better fit when a recording will be listened to repeatedly or published as finished media. Audiobooks, branded voiceovers, lessons, narrated videos, and character performances can benefit from its emphasis on expressive speech and vocal consistency.

For live assistants, rapid dialogue, or applications where users expect nearly immediate responses, the faster Speech-02-Turbo model or a newer low-latency speech offering may be a better choice. That does not mean Speech-02-HD is unusable for streaming; streaming is documented for the family. It means that the HD model's quality-oriented positioning may not align with a workload where response time is the primary metric.

Editorially, its cost profile is difficult to compare directly because current first-party pricing for this exact model is not verified. Alibaba Cloud Model Studio documentation lists a China-region price of CNY 3.5 per 10,000 characters. This is a partner-hosted price, not a confirmed universal MiniMax price, and regional taxes, billing rules, access conditions, or platform markups may differ. MiniMax's current pricing materials emphasize Speech 2.8-HD rather than Speech-02-HD, so users should verify the active endpoint and rate before planning production costs.

Where it fits in MiniMax's current catalog

Speech-02-HD is best treated as an older or legacy-generation model rather than the central HD speech offering in MiniMax's current global catalog. MiniMax's current public pricing page highlights Speech 2.8-HD, and newer speech generations have become more prominent. Some partner-hosted services may continue to expose Speech-02-HD even if it is no longer listed as a core model on MiniMax's main pricing page.

This distinction matters for deployment planning. A model can remain technically accessible through a partner while having less predictable long-term availability, documentation, or pricing than a current first-party model. Before integrating Speech-02-HD into a durable production workflow, check the exact provider, region, endpoint name, rate limits, and lifecycle notice.

Best use cases

  • Audiobook and long-form narration: expressive delivery and multiple audio formats suit published spoken content.
  • Professional voiceovers: pitch, speed, volume, and voice controls can help adapt delivery to different projects.
  • Multilingual education: support for many reported languages can assist with lessons, language materials, and instructional audio.
  • Digital characters: voice cloning and emotional control can support personalized characters and interactive media.
  • Advertising and branded audio: controllable delivery is useful for short promotional scripts and campaign variations.
  • Voice prototyping: zero-shot cloning can help test a character or narrator concept before committing to a longer voice-production process.

When to choose Speech-02-HD

Choose Speech-02-HD when the priority is polished, expressive multilingual speech and the workflow can tolerate a quality-focused generation profile. It is particularly suitable when you need voice cloning without building a speaker-specific training pipeline, or when the resulting audio will be reused in a recording, lesson, audiobook, advertisement, or character experience.

Choose another option when the application needs live, extremely low-latency conversation; general text reasoning; coding; tool use; web access; structured data generation; or image and video output. Speech-02-HD cannot replace a language model or an agent model in those workflows. A newer MiniMax speech model may also be preferable when current first-party support, clearer lifecycle status, or access to the latest speech capabilities is more important than compatibility with this older model.

Limitations and points to verify

The most important limitation is lifecycle uncertainty. Speech-02-HD is not emphasized on MiniMax's current core pricing catalog, and direct current pricing for the exact model is not fully verified. Availability may depend on a partner platform such as Alibaba Cloud Model Studio or on regional access.

The model also has a narrow modality profile: it accepts text and can use reference audio for cloning, but its output is speech audio. It is not a reasoning model, does not provide coding capability, and does not support tools, function calling, structured JSON, embeddings, image generation, or video generation. The maximum output length has not been published in the supplied sources, so applications should not assume an unlimited narration length and should test how long scripts are handled by the selected endpoint.

Finally, provider-reported language coverage and expressive controls describe intended or documented capabilities, not a guarantee that every voice, language, accent, or emotional instruction will perform equally well. Production users should test representative scripts, pronunciation, cloned-voice permissions, audio formats, latency, and per-character cost before switching from evaluation to deployment.


Answers to Frequently Asked Questions

Is MiniMax Speech-02-HD still a current production model, and what does it cost?
Speech-02-HD may be considered an older or legacy-generation model because MiniMax’s current catalog emphasizes newer offerings such as Speech 2.8-HD. Alibaba Cloud Model Studio lists a China-region price of CNY 3.5 per 10,000 characters, but this is a partner-hosted rate rather than a confirmed universal MiniMax price. Users should verify the endpoint, region, availability, and current pricing before deployment.
What is the difference between Speech-02-HD and Speech-02-Turbo?
Speech-02-HD prioritizes vocal fidelity, expressive delivery, pronunciation, and natural rhythm, while Speech-02-Turbo is positioned for faster generation and more interactive use. HD is generally better for finished recordings, whereas Turbo may be more suitable for low-latency conversations.
Which languages and audio formats does Speech-02-HD support?
MiniMax reports support for 32 languages, including English, Chinese, Japanese, Korean, Spanish, French, Arabic, German, Portuguese, Hindi, and others. The Speech-02 family supports FLAC, WAV, MP3, and PCM audio output.
What is MiniMax Speech-02-HD used for?
MiniMax Speech-02-HD is a quality-focused text-to-speech model for generating expressive spoken audio. It is suitable for audiobooks, narration, voiceovers, education, advertising, digital characters, and other audio-production tasks.
Does MiniMax Speech-02-HD support voice cloning?
Yes. Speech-02-HD supports zero-shot voice cloning, allowing a short reference recording to guide the vocal identity of generated speech without conventional speaker-specific training. Cloning should only be used with the speaker’s permission.


Sources 5
Provider

About MiniMax