Speech 2.8

speech-2.8-turbo

by MiniMax · Current and accessible through the MiniMax Open Platform API

A speed-focused MiniMax text-to-speech model for real-time and interactive voice generation. Speech-2.8-Turbo supports multilingual expressive speech, configurable voices and audio settings, streaming WebSocket synthesis, and voice-cloning workflows. It is priced at $60 per 1 million generated characters, while context and maximum output limits are not documented in the supplied research.

Speech Reasoning Coding
MiniMax Speech-2.8-Turbo is the speed-focused variant of the Speech 2.8 text-to-audio family. It is designed for applications that need generated speech quickly, such as voice assistants, conversational agents, interactive characters, games, and live narration. The model accepts text and returns synthesized audio, with official documentation describing synchronous WebSocket streaming and controls for voice, speed, pitch, volume, language behavior, pronunciation, and output format.
Outputs

What speech-2.8-turbo can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Speech 2.8
Model type Other
Release date 2026-01-23
Status Current and accessible through the MiniMax Open Platform API
Knowledge cutoff notes

A knowledge cutoff is not applicable or publicly documented for this text-to-speech model.

Model notes

The exact API model identifier is speech-2.8-turbo. It is the Turbo variant of the MiniMax Speech 2.8 text-to-audio family and is positioned for lower latency than the HD variant. The model accepts text and returns synthesized speech audio. Official documentation demonstrates synchronous WebSocket streaming through the t2a_v2 endpoint, with configurable voice ID, speed, volume, pitch, language boost, pronunciation controls, sample rate, bitrate, format and channel count. Documented example output uses MP3 at 32 kHz, 128 kbps and one channel. MiniMax describes the Speech 2.8 family as supporting native sound tags, expressive speech, multilingual synthesis and high-fidelity voice cloning. The model is billed by generated characters rather than input or output tokens. Pricing may vary by service plan or resource package.

Cost

Model pricing

Input $60 per 1 million characters
Model guide

MiniMax Speech-2.8-Turbo: Low-Latency Expressive Text-to-Speech

MiniMax Speech-2.8-Turbo is a low-latency text-to-speech model for real-time voice generation. It converts text into expressive multilingual speech and supports streaming synthesis, configurable audio output, voice cloning workflows, native sound tags, and adjustable delivery controls through the MiniMax Open Platform API.

What is MiniMax Speech-2.8-Turbo?

MiniMax Speech-2.8-Turbo is a text-to-speech model from MiniMax. Its job is specific: it takes written text as input and produces spoken audio. It is not a general-purpose language model, transcription system, image model, or coding assistant. The model is intended for situations where an application needs a voice response rather than a text response.

The “Turbo” name reflects its positioning within the Speech 2.8 family. MiniMax describes this version as the lower-latency option compared with the Speech 2.8 HD variant. That makes it a practical choice for interactive experiences in which waiting for a longer rendering process would make the conversation or user interface feel slow.

The model was released on January 23, 2026, according to the supplied model information, and is currently accessible through the MiniMax Open Platform API. Its exact API model identifier is speech-2.8-turbo.

Where it fits in the MiniMax lineup

Speech-2.8-Turbo belongs to MiniMax’s speech and audio generation offerings rather than its general text, coding, image, or video model lines. MiniMax’s wider ecosystem includes separate products for coding, agents, design, video, audio, and other interactive experiences. This model is the API-oriented speech synthesis component for developers who want to add generated voices to their own software or workflows.

Within the Speech 2.8 family, Turbo is positioned around responsiveness. The HD sibling is the more relevant comparison when an application prioritizes the family’s higher-fidelity positioning over the shortest possible response time, although the supplied research does not provide a detailed benchmark or a direct price comparison between the two variants. The important distinction supported by the available information is that Turbo is intended to reduce latency.

Core capabilities and supported output

Speech-2.8-Turbo accepts text input and produces audio output. It does not accept images, audio, or video as model inputs according to the supplied specifications. The output is speech audio rather than text, images, video, music, embeddings, or structured data.

MiniMax’s Speech 2.8 materials describe several capabilities relevant to expressive synthesis:

  • Multilingual speech synthesis: the model can generate speech in multiple languages, with a language boost setting available in the documented API controls.
  • Expressive delivery: the Speech 2.8 family is described as supporting expressive speech rather than only flat, neutral pronunciation.
  • Native sound tags: MiniMax highlights sound tags as part of the family’s supported expressive controls. These can be useful when a voice experience needs more than ordinary sentence reading, although the supplied research does not define the full tag syntax.
  • Voice cloning workflows: the family is positioned for high-fidelity voice cloning. The exact cloning process, consent requirements, and account or endpoint requirements should be checked in the current MiniMax documentation before production use.
  • Streaming synthesis: official API documentation demonstrates synchronous WebSocket streaming through the t2a_v2 endpoint, allowing an application to receive generated audio progressively rather than waiting for one complete file.

These are provider-documented capabilities or positioning claims. They should not be interpreted as a guarantee that every language, voice, sound tag, or cloning workflow has identical quality or availability.

Voice and audio controls

The documented WebSocket interface provides more control than a simple text-to-speech request. Developers can configure the voice ID and adjust how the generated voice is delivered. Available controls described in the supplied research include speed, volume, pitch, language boost, and pronunciation settings.

The API also documents output parameters including sample rate, bitrate, format, and channel count. An example response uses MP3 audio at 32 kHz, 128 kbps, and one channel. This is an example configuration rather than a statement that all requests must use those values. The ability to select audio characteristics is useful when generated speech must fit a mobile application, game engine, voice pipeline, or storage budget.

For a beginner, the practical workflow is straightforward: choose a supported voice, provide the text to be spoken, select the desired language and delivery settings, then consume the returned audio stream. More advanced users can use pronunciation controls and audio parameters to make the result fit a particular product or playback environment.

Speed and cost trade-offs

Speech-2.8-Turbo’s primary practical advantage is its low-latency positioning. Streaming and the Turbo designation make it better suited to interactive responses than a speech model intended mainly for offline rendering. Examples include an assistant speaking after each user turn, a game character responding during play, or an interactive story generating short spoken passages on demand.

The trade-off is that the supplied research does not establish that Turbo delivers the highest possible fidelity in every scenario. MiniMax positions Speech 2.8 HD as the related higher-fidelity option, while Turbo is aimed at responsiveness. If the audio will be produced in advance for narration, a premium voiceover, or a fixed media asset, a slower or higher-fidelity option may deserve evaluation instead.

MiniMax lists the price for Speech-2.8-Turbo as $60 per 1 million generated characters on the cited Token Plan pricing information. Billing is based on generated characters rather than input or output tokens. The price may vary by service plan or resource package, so developers should confirm the applicable rate and quota before estimating production costs. Character-based pricing also means that long prompts, repeated narration, and verbose assistant responses directly increase usage.

Technical specifications and unavailable limits

SpecificationDocumented information
ProviderMiniMax
Model identifierspeech-2.8-turbo
Model familySpeech 2.8
Primary taskText-to-speech generation
InputText
OutputSpeech audio
StreamingYes; documented through synchronous WebSocket streaming
API endpointt2a_v2 WebSocket endpoint
Price$60 per 1 million generated characters, subject to plan or resource-package conditions
Context lengthNot publicly documented in the supplied research
Maximum outputNot publicly documented in the supplied research

There is no documented context-window figure or maximum-output limit in the supplied materials. That information should not be inferred from the model’s streaming support. In practice, developers should check the current API documentation for request-length restrictions, maximum stream duration, connection limits, supported languages, voice availability, and error behavior before deploying the model.

Reasoning, coding, and tool support

Speech-2.8-Turbo is not a reasoning or coding model. It does not generate text answers, write software, browse the web, call tools, or return structured JSON as its primary output. The model’s role begins after an application has determined what should be said: it turns that text into speech.

A typical assistant architecture may therefore use a separate language model for conversation, planning, or tool calls and then send the final response text to Speech-2.8-Turbo. In that setup, Speech-2.8-Turbo handles the voice layer, while another component handles reasoning. This separation can be useful because it lets a product change its conversational model without replacing its voice-generation model.

Best use cases

Speech-2.8-Turbo is a strong fit when fast spoken responses matter more than treating speech generation as a one-time studio-rendering task. Suitable applications include:

  • Real-time voice assistants that need to begin responding quickly.
  • Conversational agents that speak after each user interaction.
  • Interactive characters in games, simulations, and entertainment products.
  • Multilingual narration for applications that generate content dynamically.
  • Voice interfaces that need configurable speed, pitch, volume, or pronunciation.
  • Prototypes and production services that need streamed audio rather than a fully rendered file before playback begins.
  • Voice-cloning workflows where the relevant permissions, identity safeguards, and MiniMax requirements have been satisfied.

The model is less appropriate for text generation, speech recognition, music creation, image or video production, or applications that need an all-in-one reasoning system. It may also be a less suitable first choice for a fixed, high-fidelity voiceover if latency is unimportant and the HD sibling or another specialized production voice system better matches the quality requirement.

When to choose Speech-2.8-Turbo

Choose MiniMax Speech-2.8-Turbo when the central requirement is responsive, expressive text-to-speech with API control and streaming. It is especially compelling when a user is waiting for the system to speak, because lower latency can improve the perceived responsiveness of an otherwise complex application.

Choose a different type of option when the requirement is not speech synthesis. A speech-recognition model is needed for converting recordings into text; a language model is needed for reasoning or response generation; and a music-generation system is needed for songs or instrumental audio. Within the MiniMax Speech 2.8 family, consider the HD variant when the project prioritizes its higher-fidelity positioning over Turbo’s speed focus, subject to confirming current pricing and availability.

The model’s limitations are equally important: no supplied context or output ceiling is available, quality may vary by language and voice, and the $60-per-million-character rate can become significant for long or highly conversational workloads. Testing representative scripts, languages, voices, and streaming conditions is more reliable than assuming that a provider-level capability claim will produce the same result in every use case.

Bottom line

MiniMax Speech-2.8-Turbo is a specialized, API-accessible speech model built for fast expressive voice generation. Its strongest differentiators are the Turbo latency focus, streamed WebSocket delivery, multilingual synthesis, configurable voice behavior, and support for the Speech 2.8 family’s expressive and cloning-oriented workflows. It should be evaluated as a voice-generation component rather than as a general AI model. For interactive products that need spoken responses quickly, it offers a clear speed-oriented option; for offline or maximum-fidelity production, another speech option may be more appropriate.


Answers to Frequently Asked Questions

What controls and audio formats does MiniMax Speech-2.8-Turbo provide?
The API provides controls for voice ID, speed, volume, pitch, language boost, and pronunciation. Developers can also configure output properties such as sample rate, bitrate, format, and channel count. An example configuration produces mono MP3 audio at 32 kHz and 128 kbps, but these values are not mandatory for every request.
How much does MiniMax Speech-2.8-Turbo cost?
MiniMax lists Speech-2.8-Turbo at $60 per 1 million generated characters under the cited Token Plan pricing information. The applicable rate may vary by service plan or resource package, so developers should confirm current pricing and quotas before estimating costs.
Does MiniMax Speech-2.8-Turbo support streaming audio?
Yes. The model supports synchronous WebSocket streaming through the documented t2a_v2 endpoint, allowing applications to receive generated audio progressively instead of waiting for a complete audio file.
What is MiniMax Speech-2.8-Turbo used for?
MiniMax Speech-2.8-Turbo is an API-based text-to-speech model that converts written text into expressive speech audio. It is designed for interactive applications such as voice assistants, conversational agents, games, simulations, and dynamically generated narration.
How does MiniMax Speech-2.8-Turbo differ from Speech 2.8 HD?
Speech-2.8-Turbo is positioned as the lower-latency option for responsive, interactive speech generation, while Speech 2.8 HD is the more relevant choice when the project prioritizes higher-fidelity positioning over the shortest response time. The supplied information does not provide a direct benchmark or price comparison.


Sources 4
Provider

About MiniMax