Speech 2.8

Speech-2.8-HD

by MiniMax · Current and available through the MiniMax API

MiniMax Speech-2.8-HD is a high-definition text-to-speech model for expressive multilingual audio. It supports sound tags, voice cloning, configurable audio settings, HTTP synthesis, and WebSocket streaming. MiniMax lists pricing at $100 per 1 million characters, making it better suited to quality-sensitive narration and voice applications than cost-minimized workloads.

Speech
MiniMax Speech-2.8-HD is the high-definition model in MiniMax's Speech 2.8 text-to-speech family. It turns written text into expressive speech for applications such as audiobooks, podcasts, advertising, games, language learning, accessibility, and voice agents. The model is designed for users who value vocal fidelity and expressive delivery, with support for sound tags, multilingual synthesis, voice cloning, configurable audio settings, HTTP requests, and WebSocket streaming.
Outputs

What Speech-2.8-HD can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Speech 2.8
Model type Other
Release date 2026-01-23
Status Current and available through the MiniMax API
Knowledge cutoff notes

MiniMax has not published a conventional knowledge cutoff for this text-to-speech model.

Model notes

The canonical API model ID is speech-2.8-hd. It is the high-definition member of MiniMax's Speech 2.8 text-to-audio family. MiniMax documents native sound tags, high-fidelity voice cloning, multilingual synthesis, configurable audio settings, HTTP generation, and WebSocket streaming. The model is billed by generated text characters rather than token input/output. MiniMax's official pricing page lists $100 per 1 million characters. Voice cloning and voice-design features are related platform capabilities and should not be interpreted as evidence that the base model is fine-tuned by users. Conventional LLM fields such as context length, maximum output tokens, knowledge cutoff, and batch processing are not specified for this speech model.

Cost

Model pricing

Input $100 per 1 million characters for text-to-audio usage
Model guide

MiniMax Speech-2.8-HD: High-Fidelity Expressive Text-to-Speech

MiniMax Speech-2.8-HD is a high-definition text-to-speech model for expressive, natural-sounding audio. It supports native sound tags, voice cloning, multilingual synthesis, configurable audio output, standard HTTP generation, and incremental WebSocket streaming. Its main trade-off is a higher character-based price than MiniMax Speech-2.8-Turbo, making it better suited to narration and audio quality-sensitive applications than cost-minimized or latency-critical workloads.

What is MiniMax Speech-2.8-HD?

MiniMax Speech-2.8-HD is a text-to-speech model provided through the MiniMax API. Its canonical API model ID is speech-2.8-hd. Given written text and voice configuration, it produces an audio file or an audio stream rather than a text response.

The model belongs to MiniMax's Speech 2.8 family and is positioned as the high-definition option. MiniMax also lists Speech-2.8-Turbo, which is intended as a lower-cost, latency-oriented alternative. This positioning makes Speech-2.8-HD most relevant when the quality and expressiveness of the generated voice matter more than achieving the lowest possible cost or response time.

MiniMax announced the Speech 2.8 family on January 23, 2026. The specifications described here are based on MiniMax's model announcement, API documentation, WebSocket documentation, and pricing materials.

Primary purpose and main strengths

Speech-2.8-HD is built for expressive speech generation rather than general-purpose language processing. It can be used to create narration, character dialogue, announcements, and other spoken content from text. Its high-definition positioning is particularly relevant to productions where listeners are likely to notice unnatural timing, pronunciation, vocal tone, or missing emotional cues.

  • Expressive delivery: MiniMax documents native sound tags for effects and delivery cues such as breaths, sighs, chuckles, laughs, and throat clearing.
  • Voice cloning: MiniMax's Speech 2.8 materials describe high-fidelity voice cloning from a short voice sample through the platform's voice-cloning features.
  • Multilingual synthesis: MiniMax reports support for more than 40 languages and over 90 dialects in its Speech 2.8 materials.
  • Cross-lingual performance: The provider highlights improved pronunciation and accent behavior for cases such as Mandarin-to-Japanese speech.
  • Output control: Applications can configure voice identity, speed, volume, pitch, pronunciation, sample rate, bitrate, format, and channel settings.
  • Streaming delivery: The model can return a completed result over HTTP or deliver audio incrementally through the WebSocket text-to-audio API.

These capabilities make the model more than a basic read-aloud engine. For example, a scripted character can include a laugh or breath cue in the text, while an interactive application can begin playing audio as chunks arrive instead of waiting for the full synthesis request to finish.

Supported input and output

The model accepts text as its primary input and produces speech audio as its output. It is not documented as a multimodal understanding model: the supplied specifications do not indicate support for image, video, or audio input. It also does not return text, images, video, embeddings, or structured tool actions.

MiniMax's HTTP examples show MP3 output configured at a 32 kHz sample rate, 128 kbps bitrate, and mono channel. These are documented example settings rather than an assertion that every application must use them. The API also returns usage character counts and audio metadata, which can help applications monitor billing and process the generated file.

The WebSocket interface is useful for interactive systems because it supports incremental audio delivery. This can reduce the time before playback begins, although the model's high-definition positioning means it should not automatically be assumed to match a lower-latency speech model in every situation.

Sound tags, voice cloning, and voice control

One of Speech-2.8-HD's distinguishing features is support for native sound tags. These tags allow a script to communicate certain vocal or nonverbal effects directly to the synthesis system. Examples documented by MiniMax include breathing, sighing, chuckling, laughing, and throat clearing. This is useful for narration and character dialogue because the script can contain expressive cues instead of relying entirely on post-production editing.

Voice cloning can help a production maintain a recognizable speaker across multiple passages or languages. However, cloning a voice does not remove the need for permission and appropriate rights management. Applications should obtain the necessary consent before using a person's voice and should make sure that generated audio is not presented deceptively.

The API exposes controls for voice identity, speed, volume, pitch, pronunciation, sample rate, bitrate, format, and channel configuration. These controls allow developers to adapt the same model to different delivery styles. A podcast may favor a natural speaking speed and high-quality stereo workflow, while a conversational application may prioritize compact mono audio and rapid playback.

API and streaming options

Speech-2.8-HD is available through MiniMax's text-to-audio API. The HTTP endpoint is suitable for conventional request-and-response workflows: an application sends text and settings, waits for synthesis, and receives generated audio and related metadata.

The WebSocket endpoint is intended for streaming use. It can send audio incrementally, making it better suited to voice agents, live interfaces, and applications where playback should begin before the complete response has been generated. Streaming changes delivery behavior, not the model's fundamental role: Speech-2.8-HD remains a text-to-speech system rather than a conversational reasoning model.

Implementations should retain the exact model ID speech-2.8-hd, handle audio encoding according to the selected settings, and monitor the returned character usage. They should also distinguish between a successful API request and usable production audio: pronunciation, pacing, emotional interpretation, and voice permissions still require application-level testing and review.

Pricing and cost trade-offs

MiniMax lists Speech-2.8-HD at $100 per 1 million characters. The supplied pricing information describes billing by text characters, not by language-model input and output tokens. For a rough budget, an application should estimate the amount of source text it will send, account for repeated generation or retries, and consider whether streaming or multiple voice versions will increase total usage.

The principal cost comparison in the same family is Speech-2.8-Turbo. MiniMax positions Turbo as the lower-cost, latency-oriented option, while Speech-2.8-HD is the quality-focused choice. The supplied research does not provide a specific Turbo price or a measured latency difference, so those values should not be inferred. In practical terms, HD is most defensible when improved audio quality can justify the higher listed character cost.

Limits and unsupported capabilities

Speech-2.8-HD is specialized. It generates speech but does not perform the functions normally associated with a general-purpose language model. The supplied documentation does not specify a conventional context-window size, maximum output-token limit, knowledge cutoff, or batch API. Because the model outputs audio rather than tokens, an output-token limit is not a meaningful documented specification here.

  • Reasoning: No reasoning capability or reasoning score is documented.
  • Coding: The model does not generate or analyze code as a primary capability.
  • Tools and functions: Native tool calling and function calling are not documented.
  • JSON and structured output: JSON mode and structured output are not documented for the speech response.
  • Input modalities: The supplied specification identifies text input, not image, audio, or video input.
  • Fine-tuning and caching: Fine-tuning and caching are not listed as supported model features.

These limitations do not make the model unsuitable for voice agents. A voice agent can combine Speech-2.8-HD with a separate language model or application backend: another component handles conversation logic and tool calls, while Speech-2.8-HD renders the final text as speech. The model itself should not be treated as the component that understands requests, searches the web, executes code, or decides which function to call.

Best use cases

Speech-2.8-HD is a strong fit for applications where expressive, polished speech is central to the user experience:

  • Audiobook and long-form narration where consistent voice quality matters.
  • Podcast intros, advertisements, and other produced audio that benefits from controllable delivery.
  • Game and animation characters using sound tags and distinctive voices.
  • Multilingual narration and language-learning content.
  • Accessibility features that require natural-sounding spoken content.
  • Voice agents that need higher-fidelity spoken responses and can tolerate the model's character-based cost.
  • Interactive experiences that can use WebSocket streaming to begin playback progressively.

For a high-volume application producing straightforward speech with minimal expressive requirements, the lower-cost Speech-2.8-Turbo option may be more appropriate. For a system that needs text reasoning, web search, code execution, image understanding, or tool use, Speech-2.8-HD should be paired with other components rather than selected as the central model.

When to choose Speech-2.8-HD

Choose Speech-2.8-HD when the output will be heard by customers, audiences, or players and vocal realism is more important than minimizing the synthesis bill. It is especially suitable when sound tags, voice identity, multilingual delivery, configurable audio encoding, or high-fidelity narration are important requirements.

Choose a different option when the main priority is the lowest cost or fastest possible speech response. Within MiniMax's documented lineup, Speech-2.8-Turbo is the relevant sibling to evaluate for that trade-off. Also choose a separate language or multimodal model when the application needs reasoning, coding, image or video understanding, tool calls, or a documented context window.

Overall, Speech-2.8-HD is best understood as a production-oriented speech renderer: it turns carefully prepared text into expressive audio and offers both file-based and streaming delivery. Its value comes from audio quality and control, while its boundaries are the same ones that define a specialized text-to-speech model.


Answers to Frequently Asked Questions

Does MiniMax Speech-2.8-HD support streaming?
Yes. Speech-2.8-HD supports conventional HTTP requests that return completed audio and WebSocket text-to-audio streaming that delivers audio incrementally. WebSocket streaming is useful for voice agents and interactive applications that need playback to begin before the full response is generated.
How much does MiniMax Speech-2.8-HD cost?
MiniMax lists Speech-2.8-HD at $100 per 1 million characters. Billing is based on text characters, so applications should account for retries, repeated generations, and multiple voice versions when estimating costs.
What is the API model ID for MiniMax Speech-2.8-HD?
The canonical MiniMax API model ID is speech-2.8-hd.
Does MiniMax Speech-2.8-HD support voice cloning and expressive sound effects?
Yes. MiniMax documents high-fidelity voice cloning features and native sound tags for effects and delivery cues such as breaths, sighs, chuckles, laughs, and throat clearing. Voice cloning should only be used with appropriate consent and rights.
What is MiniMax Speech-2.8-HD used for?
MiniMax Speech-2.8-HD is a high-definition text-to-speech model used to turn written text into expressive audio. It is suitable for narration, podcasts, advertisements, character dialogue, multilingual content, accessibility features, and high-fidelity voice agents.


Sources 5
Provider

About MiniMax