GPT-4o Mini

GPT-4o Mini TTS

by OpenAI · Current; the canonical alias currently points to the gpt-4o-mini-tts-2025-12-15 snapshot.

OpenAI's GPT-4o Mini TTS converts text into expressive spoken audio through the Audio API. It supports delivery instructions, preset voices, multiple formats, streaming, and pricing based on text input and audio output tokens.

Speech Reasoning Coding
GPT-4o Mini TTS is OpenAI's current model for generating speech from text. Released on March 20, 2025, it is designed for fast, relatively low-cost voice generation in applications such as narration, customer service, accessibility tools, voice interfaces, and realtime audio experiences. Unlike a general-purpose GPT model, it accepts text and returns spoken audio rather than chat responses, images, code, or tool calls.
Outputs

What GPT-4o Mini TTS can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o Mini
Model type Other
Context window 2K tokens
Release date 2025-03-20
Status Current; the canonical alias currently points to the gpt-4o-mini-tts-2025-12-15 snapshot.
Knowledge cutoff notes

OpenAI does not publish a separate knowledge-cutoff date for GPT-4o Mini TTS. It is a specialized text-to-speech model, so the language-model knowledge-cutoff field is not directly applicable.

Model notes

GPT-4o Mini TTS is a specialized speech-generation model built on GPT-4o Mini and released on March 20, 2025. The model accepts text only and returns spoken audio only. Instructions can control accent, emotional range, intonation, impressions, speech speed, tone, and whispering. OpenAI lists preset voices and supports MP3, Opus, AAC, FLAC, WAV, and PCM output. The speech endpoint supports streaming output. The model documentation lists a maximum of 2,000 input tokens, while the speech endpoint reference separately lists a maximum input length of 4,096 characters. The current alias has snapshots dated 2025-03-20 and 2025-12-15, with the alias currently mapped to the later snapshot. OpenAI requires clear disclosure that generated voices are AI-generated.

Cost

Model pricing

Input $0.60 per 1M text input tokens
Output $12.00 per 1M audio output tokens
Model guide

GPT-4o Mini TTS: Features, Pricing, Voices and API Support

GPT-4o Mini TTS is OpenAI's specialized text-to-speech model for turning written text into natural-sounding spoken audio. It supports controllable delivery instructions, preset voices, several audio formats, and streaming through the Audio API, with pricing based on text input and audio output tokens.

What is GPT-4o Mini TTS?

GPT-4o Mini TTS is a specialized text-to-speech model from OpenAI. Text-to-speech, often abbreviated as TTS, means converting written words into an audio recording that sounds like spoken language. The model is accessed through OpenAI's Audio API and is intended for developers building products that need generated speech.

The canonical model ID is gpt-4o-mini-tts. OpenAI lists dated snapshots including gpt-4o-mini-tts-2025-03-20 and gpt-4o-mini-tts-2025-12-15; the current alias is documented as pointing to the later snapshot. It belongs to the GPT-4o Mini family but should not be confused with GPT-4o Mini Audio, GPT-4o Mini Realtime, or the older TTS-1 and TTS-1 HD models.

OpenAI released GPT-4o Mini TTS on March 20, 2025. The model's role in the catalog is narrow: it is a speech-generation component rather than a general conversational, reasoning, transcription, image, or video model.

How GPT-4o Mini TTS works

An application sends text to the Audio API speech endpoint and receives an audio response. The request can also include instructions describing how the voice should deliver the text. Supported directions include accent, emotional range, intonation, impressions, speaking speed, tone, and whispering.

These instructions make the model useful when the same written content needs different performances. For example, a narration application could request a calm and measured delivery, while an educational voice interface could ask for a clear, slightly slower tone. The instructions control delivery characteristics; they do not turn the service into unrestricted voice cloning.

The model uses preset artificial voices. OpenAI's current text-to-speech documentation lists voices such as alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar. OpenAI recommends marin and cedar for the best quality, although voice availability can depend on the model and current API behavior.

Input, output and supported modalities

GPT-4o Mini TTS accepts text input and produces spoken audio output. It does not accept images, audio recordings, or video as input, and it does not natively return text, images, video, embeddings, or executable actions.

CapabilityGPT-4o Mini TTS support
Text inputYes
Audio outputYes
Image, audio or video inputNo
Text outputNo
Image or video outputNo
Tool or function callingNo
StreamingYes
Fine-tuningNo

Available output formats include MP3, Opus, AAC, FLAC, WAV, and PCM. Streaming output allows an application to begin playing speech before the complete audio response has been generated. OpenAI recommends WAV or PCM when minimizing response delay is especially important.

Input limits and API details

The model documentation lists a maximum of 2,000 input tokens. The speech endpoint documentation separately specifies a maximum input length of 4,096 characters. Tokens and characters are different measurements, so developers should not treat these limits as interchangeable. In practice, an implementation should follow the stricter limit that applies to its selected API path and the service's current behavior.

The supplied documentation does not specify a maximum output-token limit for GPT-4o Mini TTS. The output is audio rather than a text completion, and the available audio formats and streaming behavior are more relevant to application design than a conventional text-generation output limit.

Because the model is specialized, it does not provide general reasoning, coding, browsing, structured-output, or tool-use capabilities. A typical architecture may therefore use another model to generate or revise the script, then send the final text to GPT-4o Mini TTS for speech generation. GPT-4o Mini TTS itself remains responsible only for turning that text into audio.

GPT-4o Mini TTS pricing

OpenAI's listed price is $0.60 per 1 million text input tokens and $12.00 per 1 million audio output tokens. The input charge covers the text supplied to the model, while the output charge covers the generated speech representation.

Audio-token pricing is not directly equivalent to pricing for a text-only response. Actual cost depends on both the amount of text submitted and the amount of speech produced. Applications that generate long narration, repeat generations during editing, or synthesize many customer-service responses should estimate both sides of the usage.

GPT-4o Mini TTS is best understood as a speed-and-cost-oriented speech option rather than a general model that happens to support audio. Its relatively focused feature set can be an advantage when an application needs speech generation without paying for capabilities it will not use. However, the supplied research does not provide a direct benchmark comparing its latency or voice quality with every other OpenAI speech model, so claims about superiority should be treated as editorial rather than provider-published facts.

Main strengths

  • Controllable delivery: Instructions can influence accent, emotion, intonation, impressions, speed, tone, and whispering.
  • Multiple formats: MP3, Opus, AAC, FLAC, WAV, and PCM support different playback, storage, and integration requirements.
  • Streaming: Applications can start playback before the complete response has finished generating.
  • Preset voice selection: Developers can select from a range of artificial voices instead of building a voice from scratch.
  • Focused API role: The model is purpose-built for speech generation, making its behavior easier to reason about in a TTS pipeline than a general multimodal model.
  • Broad application fit: Narration, accessibility features, customer-service responses, voice interfaces, and realtime audio applications are all supported use cases.

Limitations to understand

GPT-4o Mini TTS is not a complete voice assistant on its own. It cannot hold a general text conversation, perform complex reasoning, write or execute code, transcribe incoming speech, analyze an image, search the web, or invoke external tools. If an application needs those capabilities, another model or software component must handle them before or after speech generation.

The model accepts text only. It cannot take a speaker's recording and transform that voice, accept spoken instructions directly, or generate speech based on an audio or video input. It also uses preset artificial voices rather than offering unrestricted custom voice cloning.

Input size is another practical constraint. The model documentation lists 2,000 input tokens, while the speech endpoint lists 4,096 characters. Long books, scripts, or documents may need to be divided into smaller segments. Splitting content also requires care so that sentence boundaries, pronunciation, and audio playback remain natural between segments.

OpenAI requires applications using generated voices to clearly disclose that the speech is AI-generated. This disclosure should be included in the product experience where users could reasonably mistake the voice for a human recording.

Best use cases

GPT-4o Mini TTS is a good fit when the application already has text and needs that text spoken quickly in a selected voice. Suitable examples include:

  • Reading accessibility content aloud.
  • Producing short-form narration for educational or informational material.
  • Generating spoken responses for customer-service workflows.
  • Adding voice output to an application or device interface.
  • Creating realtime or near-realtime audio experiences with streaming playback.
  • Producing multiple versions of a script with different speeds, tones, or emotional directions.

It is particularly appropriate when preset voices are acceptable and the product benefits from several output formats. WAV or PCM may be useful for low-delay playback pipelines, while compressed formats such as MP3 or Opus may be more convenient for delivery and storage.

When to choose GPT-4o Mini TTS

Choose GPT-4o Mini TTS when your main requirement is controllable text-to-speech rather than general AI interaction. It is a sensible option for a developer who has a finished text script, wants selectable voices and delivery instructions, and needs streaming or standard audio-file output through an API.

Choose a different type of option when the primary problem is not speech synthesis. A transcription model is more appropriate for converting recorded speech into text. A general-purpose language model is better suited to reasoning, coding, script generation, or tool use. A realtime speech-to-speech system may be more appropriate when the application must accept live audio and respond conversationally without a separate text-only stage.

Within OpenAI's catalog, GPT-4o Mini TTS should also be evaluated separately from older TTS models such as TTS-1 and TTS-1 HD. Those models may be relevant when maintaining an existing integration, but the supplied research does not establish a current head-to-head quality, latency, or price comparison. The practical distinction that is verified here is that GPT-4o Mini TTS supports controllable instructions, multiple formats, and streaming through the current speech API.

Bottom line

GPT-4o Mini TTS is a focused OpenAI speech-generation model with text-only input, audio-only output, preset voices, delivery controls, six documented output formats, and streaming support. Its listed price is $0.60 per 1 million input text tokens plus $12 per 1 million output audio tokens. It is strongest when an application needs affordable, programmable voice generation, but it should be treated as one component in a larger system when the product also requires reasoning, transcription, conversation, image understanding, or external actions.


Answers to Frequently Asked Questions

What are the main limitations of GPT-4o Mini TTS?
GPT-4o Mini TTS accepts text only and produces audio only. It does not transcribe speech, analyze images, perform general reasoning, browse the web, call tools, support fine-tuning, or provide unrestricted voice cloning. The model documentation lists a 2,000-token input limit, while the speech endpoint specifies a maximum input length of 4,096 characters.
Does GPT-4o Mini TTS support streaming and voice instructions?
Yes. GPT-4o Mini TTS supports streaming, allowing applications to begin playing audio before the full response is generated. Developers can also provide instructions controlling delivery characteristics such as accent, emotion, intonation, speaking speed, tone, impressions, and whispering.
How much does GPT-4o Mini TTS cost?
GPT-4o Mini TTS is listed at $0.60 per 1 million text input tokens and $12.00 per 1 million audio output tokens. Total cost depends on both the amount of text submitted and the amount of speech generated.
What is GPT-4o Mini TTS?
GPT-4o Mini TTS is OpenAI’s specialized text-to-speech model for converting written text into spoken audio through the Audio API. It accepts text input, generates audio output, and is designed for applications such as narration, accessibility features, voice interfaces, and customer-service responses.
What voices and audio formats does GPT-4o Mini TTS support?
GPT-4o Mini TTS supports preset artificial voices including alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar. OpenAI recommends marin and cedar for the best quality. Supported output formats include MP3, Opus, AAC, FLAC, WAV, and PCM.


Sources 5
Provider

About OpenAI