Speech 2.6

MiniMax Speech 2.6 Turbo

by MiniMax · Legacy or superseded; current first-party availability is unverified

A speed-focused MiniMax text-to-speech model for real-time voice agents, streaming audio, multilingual synthesis, expressive controls, structured text handling, and voice cloning. Current first-party pricing and availability are unverified.

Speech Reasoning Coding
MiniMax Speech 2.6 Turbo converts written text into expressive spoken audio for interactive applications. It is the low-latency member of the MiniMax Speech 2.6 family, supporting streaming synthesis, voice cloning, multilingual speech, and controls for pronunciation and delivery. Its main trade-off is focus: it is an audio-generation model rather than a general-purpose language model, and current first-party pricing and availability for this specific variant are not verified.
Outputs

What MiniMax Speech 2.6 Turbo can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
Specifications

Technical details

Model family Speech 2.6
Model type Other
Release date 2025-10-30
Status Legacy or superseded; current first-party availability is unverified
Knowledge cutoff notes

A knowledge cutoff is not applicable or publicly documented for this text-to-speech generation model.

Model notes

Canonical API identifier is generally written as speech-2.6-turbo. MiniMax introduced Speech 2.6 on October 30, 2025 and described the family as optimized for Voice Agent scenarios, with end-to-end latency below 250 milliseconds. The Turbo variant is the low-latency member of the Speech 2.6 family. Reported capabilities include expressive speech, voice cloning, more than 40 languages, streaming synthesis, audio controls, and direct handling of URLs, email addresses, phone numbers, dates, monetary amounts, and similar formatted text. Current MiniMax public pages emphasize Speech 2.8 and current official pricing found during research lists Speech 2.8 rather than Speech 2.6, so exact current first-party pricing and availability for Speech 2.6 Turbo are unverified. Context length and token output limits are not applicable to this speech-generation model.

Model guide

MiniMax Speech 2.6 Turbo: Low-Latency Text-to-Speech for Voice Agents

MiniMax Speech 2.6 Turbo is a speed-focused text-to-speech model from MiniMax for real-time voice agents, conversational assistants, customer-service systems, games, and other applications that need expressive multilingual speech with low response latency.

What is MiniMax Speech 2.6 Turbo?

MiniMax Speech 2.6 Turbo is a text-to-speech model provided by MiniMax. It takes written input and generates spoken audio, with a design emphasis on responsiveness for applications where users expect an answer to begin playing quickly.

The model belongs to the Speech 2.6 family introduced by MiniMax on October 30, 2025. MiniMax positioned that family for Voice Agent scenarios and reported end-to-end latency below 250 milliseconds. The supplied research does not provide a separate measured latency figure for the Turbo variant, but identifies Turbo as the low-latency member of the family.

That positioning makes Speech 2.6 Turbo different from a general language model. It is not intended to write essays, reason through problems, generate images, or operate tools. Its job is to turn text into natural-sounding speech, potentially while the text is still being produced by another system.

Core capabilities

  • Text-to-speech generation: Converts text into spoken audio.
  • Streaming synthesis: Supports workflows in which text is submitted incrementally and audio chunks are returned before the full request is complete.
  • Expressive delivery: Provides controls related to voice, speed, volume, pitch, emotion, and pronunciation.
  • Voice cloning: Supports cloned and custom voices, subject to the applicable product rules and permissions.
  • Multilingual speech: MiniMax describes Speech 2.6 as supporting more than 40 languages, with native accents and dialects.
  • Structured text handling: The model is designed to handle items such as URLs, email addresses, phone numbers, dates, monetary values, and IP addresses without requiring the same degree of manual preprocessing as a traditional text-to-speech pipeline.

These features are particularly useful when the audio must sound conversational rather than like a static narration track. For example, a voice agent may need to read an order number, a web address, or a currency amount while also maintaining an appropriate speaking style.

Voice, language, and customization

Speech 2.6 supports a broad system voice library as well as cloned voices. Voice cloning can help a business maintain a consistent character or brand voice across interactions, while system voices may be simpler to deploy when custom recordings are unnecessary.

MiniMax also announced Fluent LoRA for improving fluency when cloning voices from recordings that contain accents or disfluent speech. The launch material describes this capability across more than 40 languages. The research does not establish that every language, voice, or customization option is available through every endpoint, so production teams should verify the exact language and voice support for their chosen integration.

Controls for pitch, volume, speed, emotion, and pronunciation allow developers to adapt speech to different contexts. A customer-service assistant might use a calm, measured delivery, while an interactive game character could require a more expressive style. These controls are practical generation parameters rather than evidence that the model independently reasons about the conversation.

API and streaming workflow

MiniMax documents access to its Speech models through HTTP and WebSocket interfaces. The WebSocket workflow uses task-start, task-continue, and task-finish events. This allows an application to send text in stages and receive audio chunks while synthesis continues.

Streaming is important for voice agents because waiting for an entire response before starting playback can make an otherwise fast system feel slow. With incremental synthesis, an application can begin speaking while later parts of the response are still being prepared. The actual perceived delay will depend on the surrounding language model, application server, network, buffering strategy, and audio playback pipeline, not only on Speech 2.6 Turbo.

Documented output formats for the Speech family include MP3, PCM, FLAC, and WAV. Exact format availability can vary by endpoint and streaming mode, so an implementation should check the current API documentation rather than assume that every format works in every request type.

Technical profile and known limits

AreaWhat is documented
ProviderMiniMax
Model typeText-to-speech and voice-generation model
Primary inputText
Primary outputSpoken audio
StreamingSupported
Voice cloningSupported
LanguagesMore than 40 claimed for the Speech 2.6 family
Context lengthNot publicly documented or applicable in the usual language-model sense
Maximum output tokensNot applicable to its primary audio-generation function
Reasoning and codingNot provided as model capabilities
Tool or function callingNot provided as a model capability

Speech 2.6 Turbo does not have a conventional context window or maximum output-token specification in the supplied documentation. Those measurements are normally used for text-generation models. For this model, practical limits are more likely to involve request size, audio duration, endpoint restrictions, account quotas, and service-specific limits, but the research does not verify numeric values for those limits.

The model accepts text and produces audio. The supplied specifications do not identify image, video, or audio input, and it should not be treated as a speech-to-text or transcription model. It also does not provide general-purpose text output, embeddings, image generation, video generation, or ordinary tool execution.

Speed, cost, and capability trade-offs

The central reason to consider Turbo is responsiveness. A low-latency speech model can make a conversational system feel more immediate than a batch-oriented synthesis service, particularly when combined with streaming. This matters for interruptions, turn-taking, customer support, interactive characters, and applications where a long silent pause harms the experience.

The trade-off is specialization. Speech 2.6 Turbo cannot replace the language model that plans an answer, retrieves information, applies business rules, or calls external tools. A typical voice-agent architecture may use one component to understand or generate text and Speech 2.6 Turbo to vocalize the result. Each additional component introduces its own latency, cost, failure modes, and privacy considerations.

Current first-party MiniMax pricing located during research emphasizes the newer Speech 2.8 family rather than Speech 2.6 Turbo. No verified current input or output price was found for this specific model. It should therefore not be selected on the assumption of a particular per-character, per-token, or subscription price. Teams should confirm whether the model is still available through their intended MiniMax endpoint and obtain current pricing before committing to production usage.

Best use cases

  • Real-time voice agents: Spoken responses can begin while the rest of a response is still being generated.
  • Customer-service automation: The model can provide natural-sounding replies and read structured details such as dates, amounts, and reference numbers.
  • Conversational assistants: It can serve as the speech layer for an assistant that already handles dialogue and reasoning elsewhere.
  • Interactive games and virtual characters: Expressive controls and cloned voices can support recurring characters.
  • Multilingual products: The reported language coverage can help localize voice experiences, subject to endpoint-specific verification.
  • Dynamic narration: Streaming and customization are useful when content is generated at request time rather than prepared in advance.

When to choose MiniMax Speech 2.6 Turbo

Choose Speech 2.6 Turbo when low response time is more important than access to general language-model capabilities, and when the application needs spoken output rather than transcription. It is a sensible candidate for a voice agent that needs streaming audio, expressive delivery, multilingual support, or a consistent custom voice.

It may be especially suitable when the application has a separate text-generation or orchestration layer and needs a dedicated speech engine. The model's handling of structured text can also reduce some of the formatting work required before synthesis.

Another speech model may be more appropriate if current first-party availability, a clearly published price, a documented service-level commitment, or a different language and voice catalog is more important than the Speech 2.6 Turbo feature set. A newer MiniMax speech model may also be preferable if it is the officially supported replacement. The available research indicates that MiniMax's current public materials emphasize Speech 2.8, but it does not provide enough information to make a detailed quality, price, or latency comparison.

Availability and status

Speech 2.6 Turbo is best treated as a legacy or superseded model whose current first-party availability is unverified. The model remains identifiable through historical MiniMax documentation and third-party endpoints, but the current MiniMax public website and pricing materials focus on Speech 2.8.

Before deployment, verify the canonical model identifier, which is generally written as speech-2.6-turbo, along with access permissions, supported output formats, quotas, regional availability, voice-cloning requirements, and current pricing. This is particularly important for a model family whose documentation and product lineup may change over time.


Answers to Frequently Asked Questions

Is MiniMax Speech 2.6 Turbo still available and what does it cost?
Its current first-party availability and pricing are unverified. MiniMax's newer public materials emphasize the Speech 2.8 family, so Speech 2.6 Turbo should be treated as a legacy or potentially superseded model. Before production use, verify the model identifier speech-2.6-turbo, endpoint access, quotas, regional availability, supported formats, voice-cloning requirements, and current pricing.
Can MiniMax Speech 2.6 Turbo replace a language model or transcription system?
No. Speech 2.6 Turbo is a specialized text-to-speech model that converts text into audio. It does not provide general reasoning, text generation, tool execution, speech-to-text transcription, or image and video capabilities. A typical voice-agent architecture uses a separate language model or orchestration layer to generate the text.
What languages and voice customization features does MiniMax Speech 2.6 Turbo support?
MiniMax describes the Speech 2.6 family as supporting more than 40 languages, including native accents and dialects. The model supports system voices, cloned and custom voices, and controls for speed, volume, pitch, emotion, and pronunciation. Exact availability may vary by endpoint, language, and integration.
What is MiniMax Speech 2.6 Turbo used for?
MiniMax Speech 2.6 Turbo is a text-to-speech model designed to generate natural-sounding spoken audio with low latency. It is suited to real-time voice agents, customer-service automation, conversational assistants, interactive characters, multilingual applications, and dynamic narration.
Does MiniMax Speech 2.6 Turbo support streaming text-to-speech?
Yes. Speech 2.6 Turbo supports streaming synthesis, allowing applications to send text incrementally and receive audio chunks before the full response is complete. This can reduce perceived delay in voice-agent interactions, although total latency also depends on the language model, network, buffering, and playback pipeline.


Sources 5
Provider

About MiniMax