Speech-02

Speech-02-Turbo

by MiniMax · Legacy or superseded; exact current first-party availability is unclear

MiniMax Speech-02-Turbo is a real-time-oriented text-to-speech model for multilingual expressive audio, streaming synthesis, configurable voices, and short-reference voice cloning. It is suited to interactive agents, games, customer support, and low-latency applications, but current pricing, limits, and availability are unclear as MiniMax promotes newer speech models.

Speech Reasoning Coding
MiniMax Speech-02-Turbo is the faster, interaction-focused model in MiniMax's Speech-02 text-to-speech family. It converts text into spoken audio for voice agents, games, customer-service systems, multilingual narration, and other applications where users should not have to wait for a complete audio file before playback begins. The model supports expressive synthesis, streaming generation, and short-reference voice cloning, but its current first-party status and pricing should be verified before using it in a new production system.
Outputs

What Speech-02-Turbo can produce

Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Speech-02
Model type Other
Release date 2025-04-02
Status Legacy or superseded; exact current first-party availability is unclear
Knowledge cutoff notes

Not applicable to this specialist text-to-speech model; no general factual knowledge cutoff is published.

Model notes

Canonical model identity is Speech-02-Turbo, also commonly written as speech-02-turbo or MiniMax Speech 02 Turbo. MiniMax introduced it on April 2, 2025 as the real-time-oriented member of the Speech-02 series. The model generates spoken audio from text and supports streaming synthesis, multilingual speech, expressive controls, and voice cloning from short reference recordings. MiniMax's current public pricing page prominently lists Speech 2.8 Turbo rather than Speech-02-Turbo, so no current official price was assigned to this exact model. Context length and maximum output-token limits are not published for this speech model. Reasoning, coding, and cost scores are editorial estimates for a specialist speech model, not vendor benchmarks.

Model guide

MiniMax Speech-02-Turbo: Low-Latency Multilingual Text-to-Speech

MiniMax Speech-02-Turbo is a real-time-oriented text-to-speech model for generating expressive multilingual speech with low latency. It supports streaming audio generation, voice cloning from short reference recordings, configurable voice controls, and common audio formats. Its main advantage is responsiveness for interactive voice applications, while its main limitations are unclear current availability, unpublished context and output limits, and the fact that newer MiniMax Speech 2.6 and Speech 2.8 models now receive greater attention.

What is MiniMax Speech-02-Turbo?

MiniMax Speech-02-Turbo is a proprietary text-to-speech model from MiniMax. Text-to-speech, commonly abbreviated as TTS, converts written text into spoken audio. Unlike a general-purpose language model, Speech-02-Turbo is not intended to write essays, answer questions, generate code, or reason over documents. Its job is to produce a voice that can be played back to a listener.

The model was introduced on April 2, 2025 as part of the Speech-02 family. MiniMax positioned the Turbo version for real-time and interactive scenarios, while Speech-02-HD was aimed more at higher-fidelity production work such as voiceovers and audiobooks. This distinction is important: Speech-02-Turbo is primarily a responsiveness choice, not a general-purpose AI assistant.

MiniMax is the provider. The model has also been identified in technical and API contexts as speech-02-turbo or MiniMax Speech 02 Turbo. However, MiniMax's newer public materials prominently feature Speech 2.6 and Speech 2.8 model lines. The exact current availability of Speech-02-Turbo is therefore not fully clear from the reviewed first-party model information.

Core capabilities and supported audio

Speech-02-Turbo takes text as its primary input and returns spoken audio. MiniMax describes the Speech-02 family as supporting more than 30 languages, with native-style accents across multiple language groups. The available language and voice behavior can depend on the active API configuration, so an application should confirm the exact language list and voice options exposed by its chosen endpoint.

The model supports expressive speech synthesis rather than producing only flat, uniform narration. API controls can include voice selection, speaking speed, volume, pitch, language selection, language boosting, and pronunciation settings. These controls are useful when the same text must sound more conversational, more formal, faster, slower, or better suited to a particular audience.

MiniMax's Speech-02 documentation lists common output formats including MP3, WAV, FLAC, and PCM. The API can also expose technical audio settings such as sample rate, bitrate, and channel count. Exact parameter availability may vary between interfaces, so production integrations should follow the documentation for the endpoint being used.

Streaming generation for interactive applications

The most important practical distinction of the Turbo variant is streaming. Through MiniMax's text-to-audio WebSocket API, an application can send text incrementally and receive audio chunks while synthesis is still underway. A WebSocket is a persistent connection that allows both sides to exchange data continuously, rather than waiting for a separate request and response for every piece of content.

Streaming can reduce perceived waiting time in a voice assistant, game character, live support tool, or conversational agent. The application can begin playback as soon as it has enough audio, while later portions continue arriving. This does not necessarily mean that the model produces higher-quality audio than a higher-fidelity alternative; it means that its workflow is better suited to situations where responsiveness matters.

Voice cloning and expressive control

The Speech-02 family supports zero-shot voice cloning from a short reference recording. MiniMax states that approximately 10 seconds of audio can be used to reproduce a speaker's vocal characteristics. Zero-shot means that the system can attempt to imitate the supplied voice without a lengthy speaker-specific training process.

Voice cloning should be used only with appropriate permission from the speaker and in compliance with applicable law and platform policies. A short recording may capture vocal identity, but it does not guarantee that every language, emotion, pronunciation, or delivery style will sound equally natural. The result can also depend on the quality of the reference recording and the settings supplied to the API.

Specifications, limits, and availability

CategoryVerified information
ProviderMiniMax
Model familySpeech-02
Release dateApril 2, 2025
Primary inputText
Primary outputSpoken audio
LanguagesMore than 30 languages claimed for the Speech-02 family
StreamingSupported through the documented text-to-audio WebSocket API
Voice cloningSupported from a short reference recording; MiniMax cites approximately 10 seconds
Context lengthNot published in the reviewed authoritative sources
Maximum output tokensNot applicable or not published for this speech model
Current statusLegacy or potentially superseded; current first-party availability is unclear

There is no verified context-window figure or maximum output-token limit for Speech-02-Turbo in the supplied sources. Those measurements are common for language models but are not necessarily the most useful way to describe a speech synthesis endpoint. In practice, an integration still needs to verify maximum text length, request limits, timeout behavior, and audio duration limits for the specific service endpoint it uses. Because those details were not established by the reviewed material, they should not be assumed.

The same caution applies to lifecycle status. MiniMax now promotes newer Speech 2.6 and Speech 2.8 lines, while the current public model overview does not clearly establish whether Speech-02-Turbo remains broadly available, restricted to an older endpoint, or fully superseded. Developers considering a new deployment should test access, review current documentation, and confirm that the model will remain available for the intended project.

Pricing and API access

No current official price was assigned to the exact Speech-02-Turbo model in the supplied research. MiniMax's current public pricing materials prominently list newer speech offerings, including Speech 2.8 Turbo, rather than clearly publishing a current price for Speech-02-Turbo. It would therefore be misleading to quote a per-character, per-minute, or per-request rate for this model.

Pricing may also depend on the product surface, billing unit, audio configuration, or account region. Before committing to a production architecture, confirm the model identifier, billing metric, quota policy, and any minimum charges in the current MiniMax documentation or account console. A newer model may be easier to price and support even if the older Turbo identity can still be accessed.

Reasoning, coding, and tool support

Speech-02-Turbo does not provide general-purpose reasoning or coding capabilities. It can synthesize speech from text, but it is not a language model for planning tasks, writing programs, analyzing files, or answering factual questions. It also does not generate images or video, create embeddings, or offer ordinary function-calling capabilities according to the supplied model assessment.

This means a voice assistant normally needs multiple components: a language model to interpret the user's request and produce a response, application logic or tools to perform actions, and a speech model such as Speech-02-Turbo to read the response aloud. Speech-02-Turbo can be the audio layer in that system, but it should not be treated as the complete conversational agent.

The model's input and output are best understood as text-to-audio rather than fully multimodal reasoning. Although the broader Speech-02 family may use reference audio for voice cloning, the model does not thereby become a general audio-understanding system. The supplied research does not verify capabilities such as speech recognition, audio transcription, video understanding, or autonomous tool execution.

Strengths and trade-offs

Speech-02-Turbo's principal strength is the combination of low-latency positioning and streaming delivery. That combination is valuable when a delay of several seconds harms the user experience. A live game character, telephone-style assistant, or customer-support voice interface can begin responding without waiting for the entire response to be rendered.

It also offers a useful range of speech controls. Multilingual output, expressive delivery, voice cloning, pronunciation settings, and common audio formats make it adaptable to more than one fixed narration workflow. These are provider-documented capabilities or claims, not independent quality guarantees: naturalness and consistency still need to be evaluated with the actual voices, languages, and scripts used by an application.

The trade-off is that a real-time-oriented model may not be the best choice when maximum production fidelity is more important than response speed. Speech-02-HD was positioned for higher-fidelity voiceover and audiobook use, while newer Speech 2.6 and 2.8 models may be more appropriate where current support, quality, or documented availability matters. The supplied evidence does not establish a universal quality ranking, so the choice should be tested with representative audio rather than inferred from model names alone.

Best use cases

  • Real-time voice agents: Streaming output can help assistants begin speaking promptly after generating a response.
  • Interactive customer support: Applications can use multilingual synthesis and configurable voices for responsive support experiences.
  • Games and virtual characters: Low-latency speech and expressive controls suit characters that must react during live interactions.
  • Multilingual narration: The model can support localized spoken content across the languages exposed by the active service.
  • Personalized audio: Voice cloning can create a consistent authorized voice for a brand, character, or individual.
  • Applications that stream results: WebSocket delivery is useful when audio should be played as it is generated instead of downloaded only after completion.

When to choose Speech-02-Turbo

Choose Speech-02-Turbo when responsiveness is a central requirement, the target language and voice behavior meet your needs, and the endpoint is still available under acceptable commercial terms. It is especially suitable for prototypes or production systems in which the first audible response is more important than maximum studio-style fidelity.

Consider another option when you need a clearly current and actively promoted MiniMax speech model, a published price for a new deployment, or higher-fidelity narration for audiobooks and voiceovers. Speech-02-HD was positioned toward that higher-fidelity role within the older family. Newer Speech 2.6 and Speech 2.8 models may be more appropriate if they offer better documented availability or a supported migration path.

Regardless of the choice, test complete user journeys rather than isolated sentences. Compare first-audio latency, total generation time, pronunciation of product names, multilingual consistency, emotional delivery, voice-cloning authorization, audio format compatibility, and behavior under the application's expected text lengths. These tests are particularly important because the public information does not provide a current price, context limit, or guaranteed lifecycle commitment for Speech-02-Turbo.

Bottom line

MiniMax Speech-02-Turbo is a specialist text-to-speech model built around fast, interactive audio generation. Its documented strengths include multilingual expressive speech, streaming synthesis, configurable voice controls, common audio formats, and short-reference voice cloning. It is not a reasoning or coding model, and it should normally be combined with other software when building a complete voice assistant.

The model's main practical concern is its uncertain current status. Because MiniMax now highlights newer speech families and does not clearly publish current pricing or lifecycle information for Speech-02-Turbo, it is best treated as a model to validate carefully rather than an automatically safe default for a new long-lived deployment.


Answers to Frequently Asked Questions

Does MiniMax Speech-02-Turbo support streaming audio?
Yes. Speech-02-Turbo supports streaming through MiniMax's text-to-audio WebSocket API, allowing applications to receive and play audio chunks while the rest of the text is still being synthesized.
What is MiniMax Speech-02-Turbo used for?
MiniMax Speech-02-Turbo is a text-to-speech model designed to convert written text into spoken audio, especially for real-time voice agents, interactive customer support, games, virtual characters, and multilingual applications.
How many languages does MiniMax Speech-02-Turbo support?
MiniMax states that the Speech-02 family supports more than 30 languages, with native-style accents across multiple language groups. The exact languages and voice options depend on the active API endpoint and configuration.
Is MiniMax Speech-02-Turbo still available and how much does it cost?
Its current availability and official pricing are unclear. MiniMax's newer public materials prominently feature Speech 2.6 and Speech 2.8, and no current official price was established for Speech-02-Turbo. Developers should verify the model endpoint, billing terms, quotas, and lifecycle status before using it in production.
Can MiniMax Speech-02-Turbo clone a voice?
Yes. The Speech-02 family supports zero-shot voice cloning from a short reference recording, with MiniMax citing approximately 10 seconds of audio. Voice cloning should be used only with the speaker's permission and in compliance with applicable laws and policies.


Sources 5
Provider

About MiniMax