What is MiniMax Speech-2.8-HD?
MiniMax Speech-2.8-HD is a text-to-speech model provided through the MiniMax API. Its canonical API model ID is speech-2.8-hd. Given written text and voice configuration, it produces an audio file or an audio stream rather than a text response.
The model belongs to MiniMax's Speech 2.8 family and is positioned as the high-definition option. MiniMax also lists Speech-2.8-Turbo, which is intended as a lower-cost, latency-oriented alternative. This positioning makes Speech-2.8-HD most relevant when the quality and expressiveness of the generated voice matter more than achieving the lowest possible cost or response time.
MiniMax announced the Speech 2.8 family on January 23, 2026. The specifications described here are based on MiniMax's model announcement, API documentation, WebSocket documentation, and pricing materials.
Primary purpose and main strengths
Speech-2.8-HD is built for expressive speech generation rather than general-purpose language processing. It can be used to create narration, character dialogue, announcements, and other spoken content from text. Its high-definition positioning is particularly relevant to productions where listeners are likely to notice unnatural timing, pronunciation, vocal tone, or missing emotional cues.
- Expressive delivery: MiniMax documents native sound tags for effects and delivery cues such as breaths, sighs, chuckles, laughs, and throat clearing.
- Voice cloning: MiniMax's Speech 2.8 materials describe high-fidelity voice cloning from a short voice sample through the platform's voice-cloning features.
- Multilingual synthesis: MiniMax reports support for more than 40 languages and over 90 dialects in its Speech 2.8 materials.
- Cross-lingual performance: The provider highlights improved pronunciation and accent behavior for cases such as Mandarin-to-Japanese speech.
- Output control: Applications can configure voice identity, speed, volume, pitch, pronunciation, sample rate, bitrate, format, and channel settings.
- Streaming delivery: The model can return a completed result over HTTP or deliver audio incrementally through the WebSocket text-to-audio API.
These capabilities make the model more than a basic read-aloud engine. For example, a scripted character can include a laugh or breath cue in the text, while an interactive application can begin playing audio as chunks arrive instead of waiting for the full synthesis request to finish.
Supported input and output
The model accepts text as its primary input and produces speech audio as its output. It is not documented as a multimodal understanding model: the supplied specifications do not indicate support for image, video, or audio input. It also does not return text, images, video, embeddings, or structured tool actions.
MiniMax's HTTP examples show MP3 output configured at a 32 kHz sample rate, 128 kbps bitrate, and mono channel. These are documented example settings rather than an assertion that every application must use them. The API also returns usage character counts and audio metadata, which can help applications monitor billing and process the generated file.
The WebSocket interface is useful for interactive systems because it supports incremental audio delivery. This can reduce the time before playback begins, although the model's high-definition positioning means it should not automatically be assumed to match a lower-latency speech model in every situation.
Sound tags, voice cloning, and voice control
One of Speech-2.8-HD's distinguishing features is support for native sound tags. These tags allow a script to communicate certain vocal or nonverbal effects directly to the synthesis system. Examples documented by MiniMax include breathing, sighing, chuckling, laughing, and throat clearing. This is useful for narration and character dialogue because the script can contain expressive cues instead of relying entirely on post-production editing.
Voice cloning can help a production maintain a recognizable speaker across multiple passages or languages. However, cloning a voice does not remove the need for permission and appropriate rights management. Applications should obtain the necessary consent before using a person's voice and should make sure that generated audio is not presented deceptively.
The API exposes controls for voice identity, speed, volume, pitch, pronunciation, sample rate, bitrate, format, and channel configuration. These controls allow developers to adapt the same model to different delivery styles. A podcast may favor a natural speaking speed and high-quality stereo workflow, while a conversational application may prioritize compact mono audio and rapid playback.
API and streaming options
Speech-2.8-HD is available through MiniMax's text-to-audio API. The HTTP endpoint is suitable for conventional request-and-response workflows: an application sends text and settings, waits for synthesis, and receives generated audio and related metadata.
The WebSocket endpoint is intended for streaming use. It can send audio incrementally, making it better suited to voice agents, live interfaces, and applications where playback should begin before the complete response has been generated. Streaming changes delivery behavior, not the model's fundamental role: Speech-2.8-HD remains a text-to-speech system rather than a conversational reasoning model.
Implementations should retain the exact model ID speech-2.8-hd, handle audio encoding according to the selected settings, and monitor the returned character usage. They should also distinguish between a successful API request and usable production audio: pronunciation, pacing, emotional interpretation, and voice permissions still require application-level testing and review.
Pricing and cost trade-offs
MiniMax lists Speech-2.8-HD at $100 per 1 million characters. The supplied pricing information describes billing by text characters, not by language-model input and output tokens. For a rough budget, an application should estimate the amount of source text it will send, account for repeated generation or retries, and consider whether streaming or multiple voice versions will increase total usage.
The principal cost comparison in the same family is Speech-2.8-Turbo. MiniMax positions Turbo as the lower-cost, latency-oriented option, while Speech-2.8-HD is the quality-focused choice. The supplied research does not provide a specific Turbo price or a measured latency difference, so those values should not be inferred. In practical terms, HD is most defensible when improved audio quality can justify the higher listed character cost.
Limits and unsupported capabilities
Speech-2.8-HD is specialized. It generates speech but does not perform the functions normally associated with a general-purpose language model. The supplied documentation does not specify a conventional context-window size, maximum output-token limit, knowledge cutoff, or batch API. Because the model outputs audio rather than tokens, an output-token limit is not a meaningful documented specification here.
- Reasoning: No reasoning capability or reasoning score is documented.
- Coding: The model does not generate or analyze code as a primary capability.
- Tools and functions: Native tool calling and function calling are not documented.
- JSON and structured output: JSON mode and structured output are not documented for the speech response.
- Input modalities: The supplied specification identifies text input, not image, audio, or video input.
- Fine-tuning and caching: Fine-tuning and caching are not listed as supported model features.
These limitations do not make the model unsuitable for voice agents. A voice agent can combine Speech-2.8-HD with a separate language model or application backend: another component handles conversation logic and tool calls, while Speech-2.8-HD renders the final text as speech. The model itself should not be treated as the component that understands requests, searches the web, executes code, or decides which function to call.
Best use cases
Speech-2.8-HD is a strong fit for applications where expressive, polished speech is central to the user experience:
- Audiobook and long-form narration where consistent voice quality matters.
- Podcast intros, advertisements, and other produced audio that benefits from controllable delivery.
- Game and animation characters using sound tags and distinctive voices.
- Multilingual narration and language-learning content.
- Accessibility features that require natural-sounding spoken content.
- Voice agents that need higher-fidelity spoken responses and can tolerate the model's character-based cost.
- Interactive experiences that can use WebSocket streaming to begin playback progressively.
For a high-volume application producing straightforward speech with minimal expressive requirements, the lower-cost Speech-2.8-Turbo option may be more appropriate. For a system that needs text reasoning, web search, code execution, image understanding, or tool use, Speech-2.8-HD should be paired with other components rather than selected as the central model.
When to choose Speech-2.8-HD
Choose Speech-2.8-HD when the output will be heard by customers, audiences, or players and vocal realism is more important than minimizing the synthesis bill. It is especially suitable when sound tags, voice identity, multilingual delivery, configurable audio encoding, or high-fidelity narration are important requirements.
Choose a different option when the main priority is the lowest cost or fastest possible speech response. Within MiniMax's documented lineup, Speech-2.8-Turbo is the relevant sibling to evaluate for that trade-off. Also choose a separate language or multimodal model when the application needs reasoning, coding, image or video understanding, tool calls, or a documented context window.
Overall, Speech-2.8-HD is best understood as a production-oriented speech renderer: it turns carefully prepared text into expressive audio and offers both file-based and streaming delivery. Its value comes from audio quality and control, while its boundaries are the same ones that define a specialized text-to-speech model.

