What is GPT-4o Mini TTS?
GPT-4o Mini TTS is a specialized text-to-speech model from OpenAI. Text-to-speech, often abbreviated as TTS, means converting written words into an audio recording that sounds like spoken language. The model is accessed through OpenAI's Audio API and is intended for developers building products that need generated speech.
The canonical model ID is gpt-4o-mini-tts. OpenAI lists dated snapshots including gpt-4o-mini-tts-2025-03-20 and gpt-4o-mini-tts-2025-12-15; the current alias is documented as pointing to the later snapshot. It belongs to the GPT-4o Mini family but should not be confused with GPT-4o Mini Audio, GPT-4o Mini Realtime, or the older TTS-1 and TTS-1 HD models.
OpenAI released GPT-4o Mini TTS on March 20, 2025. The model's role in the catalog is narrow: it is a speech-generation component rather than a general conversational, reasoning, transcription, image, or video model.
How GPT-4o Mini TTS works
An application sends text to the Audio API speech endpoint and receives an audio response. The request can also include instructions describing how the voice should deliver the text. Supported directions include accent, emotional range, intonation, impressions, speaking speed, tone, and whispering.
These instructions make the model useful when the same written content needs different performances. For example, a narration application could request a calm and measured delivery, while an educational voice interface could ask for a clear, slightly slower tone. The instructions control delivery characteristics; they do not turn the service into unrestricted voice cloning.
The model uses preset artificial voices. OpenAI's current text-to-speech documentation lists voices such as alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar. OpenAI recommends marin and cedar for the best quality, although voice availability can depend on the model and current API behavior.
Input, output and supported modalities
GPT-4o Mini TTS accepts text input and produces spoken audio output. It does not accept images, audio recordings, or video as input, and it does not natively return text, images, video, embeddings, or executable actions.
| Capability | GPT-4o Mini TTS support |
|---|---|
| Text input | Yes |
| Audio output | Yes |
| Image, audio or video input | No |
| Text output | No |
| Image or video output | No |
| Tool or function calling | No |
| Streaming | Yes |
| Fine-tuning | No |
Available output formats include MP3, Opus, AAC, FLAC, WAV, and PCM. Streaming output allows an application to begin playing speech before the complete audio response has been generated. OpenAI recommends WAV or PCM when minimizing response delay is especially important.
Input limits and API details
The model documentation lists a maximum of 2,000 input tokens. The speech endpoint documentation separately specifies a maximum input length of 4,096 characters. Tokens and characters are different measurements, so developers should not treat these limits as interchangeable. In practice, an implementation should follow the stricter limit that applies to its selected API path and the service's current behavior.
The supplied documentation does not specify a maximum output-token limit for GPT-4o Mini TTS. The output is audio rather than a text completion, and the available audio formats and streaming behavior are more relevant to application design than a conventional text-generation output limit.
Because the model is specialized, it does not provide general reasoning, coding, browsing, structured-output, or tool-use capabilities. A typical architecture may therefore use another model to generate or revise the script, then send the final text to GPT-4o Mini TTS for speech generation. GPT-4o Mini TTS itself remains responsible only for turning that text into audio.
GPT-4o Mini TTS pricing
OpenAI's listed price is $0.60 per 1 million text input tokens and $12.00 per 1 million audio output tokens. The input charge covers the text supplied to the model, while the output charge covers the generated speech representation.
Audio-token pricing is not directly equivalent to pricing for a text-only response. Actual cost depends on both the amount of text submitted and the amount of speech produced. Applications that generate long narration, repeat generations during editing, or synthesize many customer-service responses should estimate both sides of the usage.
GPT-4o Mini TTS is best understood as a speed-and-cost-oriented speech option rather than a general model that happens to support audio. Its relatively focused feature set can be an advantage when an application needs speech generation without paying for capabilities it will not use. However, the supplied research does not provide a direct benchmark comparing its latency or voice quality with every other OpenAI speech model, so claims about superiority should be treated as editorial rather than provider-published facts.
Main strengths
- Controllable delivery: Instructions can influence accent, emotion, intonation, impressions, speed, tone, and whispering.
- Multiple formats: MP3, Opus, AAC, FLAC, WAV, and PCM support different playback, storage, and integration requirements.
- Streaming: Applications can start playback before the complete response has finished generating.
- Preset voice selection: Developers can select from a range of artificial voices instead of building a voice from scratch.
- Focused API role: The model is purpose-built for speech generation, making its behavior easier to reason about in a TTS pipeline than a general multimodal model.
- Broad application fit: Narration, accessibility features, customer-service responses, voice interfaces, and realtime audio applications are all supported use cases.
Limitations to understand
GPT-4o Mini TTS is not a complete voice assistant on its own. It cannot hold a general text conversation, perform complex reasoning, write or execute code, transcribe incoming speech, analyze an image, search the web, or invoke external tools. If an application needs those capabilities, another model or software component must handle them before or after speech generation.
The model accepts text only. It cannot take a speaker's recording and transform that voice, accept spoken instructions directly, or generate speech based on an audio or video input. It also uses preset artificial voices rather than offering unrestricted custom voice cloning.
Input size is another practical constraint. The model documentation lists 2,000 input tokens, while the speech endpoint lists 4,096 characters. Long books, scripts, or documents may need to be divided into smaller segments. Splitting content also requires care so that sentence boundaries, pronunciation, and audio playback remain natural between segments.
OpenAI requires applications using generated voices to clearly disclose that the speech is AI-generated. This disclosure should be included in the product experience where users could reasonably mistake the voice for a human recording.
Best use cases
GPT-4o Mini TTS is a good fit when the application already has text and needs that text spoken quickly in a selected voice. Suitable examples include:
- Reading accessibility content aloud.
- Producing short-form narration for educational or informational material.
- Generating spoken responses for customer-service workflows.
- Adding voice output to an application or device interface.
- Creating realtime or near-realtime audio experiences with streaming playback.
- Producing multiple versions of a script with different speeds, tones, or emotional directions.
It is particularly appropriate when preset voices are acceptable and the product benefits from several output formats. WAV or PCM may be useful for low-delay playback pipelines, while compressed formats such as MP3 or Opus may be more convenient for delivery and storage.
When to choose GPT-4o Mini TTS
Choose GPT-4o Mini TTS when your main requirement is controllable text-to-speech rather than general AI interaction. It is a sensible option for a developer who has a finished text script, wants selectable voices and delivery instructions, and needs streaming or standard audio-file output through an API.
Choose a different type of option when the primary problem is not speech synthesis. A transcription model is more appropriate for converting recorded speech into text. A general-purpose language model is better suited to reasoning, coding, script generation, or tool use. A realtime speech-to-speech system may be more appropriate when the application must accept live audio and respond conversationally without a separate text-only stage.
Within OpenAI's catalog, GPT-4o Mini TTS should also be evaluated separately from older TTS models such as TTS-1 and TTS-1 HD. Those models may be relevant when maintaining an existing integration, but the supplied research does not establish a current head-to-head quality, latency, or price comparison. The practical distinction that is verified here is that GPT-4o Mini TTS supports controllable instructions, multiple formats, and streaming through the current speech API.
Bottom line
GPT-4o Mini TTS is a focused OpenAI speech-generation model with text-only input, audio-only output, preset voices, delivery controls, six documented output formats, and streaming support. Its listed price is $0.60 per 1 million input text tokens plus $12 per 1 million output audio tokens. It is strongest when an application needs affordable, programmable voice generation, but it should be treated as one component in a larger system when the product also requires reasoning, transcription, conversation, image understanding, or external actions.

