TTS-1

TTS-1 HD

by OpenAI · Current; available through the OpenAI Audio API speech endpoint

OpenAI’s TTS-1 HD converts text into generated speech through the Audio API. It prioritizes quality over the lowest latency, supports preset voices, adjustable playback speed, audio streaming, and MP3, Opus, AAC, FLAC, WAV, and PCM output. Requests can contain up to 4,096 characters, and pricing is $30 per 1 million characters.

Speech Reasoning Coding
TTS-1 HD is OpenAI’s text-to-speech model for applications that need higher-quality generated speech rather than the lowest possible response latency. It accepts written text through the Audio API and returns speech audio in several downloadable formats. The model is suited to narration, accessibility, announcements, educational content, and voice interfaces where sound quality matters more than minimizing generation time.
Outputs

What TTS-1 HD can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family TTS-1
Model type Other
Release date 2023-11-06
Status Current; available through the OpenAI Audio API speech endpoint
Knowledge cutoff notes

OpenAI does not publish a knowledge-cutoff date for TTS-1 HD. As a text-to-speech model, its primary function is speech generation rather than knowledge-based text reasoning.

Model notes

The canonical API model ID is tts-1-hd. OpenAI introduced it on November 6, 2023, alongside TTS-1. It is optimized for quality, while TTS-1 prioritizes lower latency. The speech endpoint accepts a maximum of 4,096 characters per request and supports MP3, Opus, AAC, FLAC, WAV, and PCM output. Supported voice selection depends on the model. The instructions parameter does not work with TTS-1 HD. The endpoint supports audio streaming, but SSE streaming is not supported for tts-1 or tts-1-hd. OpenAI’s current guide recommends GPT-4o mini TTS for intelligent real-time applications but still lists TTS-1 HD as an available quality-oriented speech model.

Cost

Model pricing

Input $30 per 1 million characters
Output Included in speech-generation pricing; output is billed by input characters rather than output tokens
Model guide

TTS-1 HD: OpenAI’s High-Quality Text-to-Speech Model

TTS-1 HD is OpenAI’s quality-oriented text-to-speech model. It converts text into natural-sounding spoken audio through the Audio API speech endpoint, supporting formats such as MP3, Opus, AAC, FLAC, WAV, and PCM. Compared with TTS-1, it favors speech quality over the lowest possible latency. The model accepts up to 4,096 characters per request, supports preset voices and playback-speed control, and is priced at $30 per 1 million characters.

What is TTS-1 HD?

TTS-1 HD is OpenAI’s quality-focused text-to-speech model. Text-to-speech, often abbreviated as TTS, converts written words into spoken audio. TTS-1 HD performs this conversion through OpenAI’s Audio API speech endpoint using the model identifier tts-1-hd.

The model is positioned as the higher-quality counterpart to TTS-1. The distinction is mainly about the trade-off between audio quality and latency: TTS-1 is designed for lower-latency speech generation, while TTS-1 HD is intended for cases where a more quality-oriented result is worth accepting potentially higher latency or cost.

TTS-1 HD is not a general-purpose language model. Its job is to generate speech from supplied text. It is not designed for reasoning, coding, transcription, image analysis, web search, or general conversation by itself.

Where TTS-1 HD fits in OpenAI’s catalog

TTS-1 HD is part of OpenAI’s Audio API model lineup and is currently documented as an available quality-oriented speech model. OpenAI introduced TTS-1 HD on November 6, 2023, alongside TTS-1. The current model catalog uses tts-1-hd as its canonical API identifier.

OpenAI’s newer text-to-speech documentation recommends GPT-4o mini TTS for intelligent real-time applications, but it continues to document TTS-1 HD as a speech-generation option. In practical terms, TTS-1 HD remains relevant when an application already has text and needs downloadable or streamed speech, particularly when quality is more important than the minimum possible delay.

Inputs, outputs, and supported modalities

TTS-1 HD has a narrow, clearly defined modality profile:

  • Input: Text.
  • Output: Generated speech audio.
  • Input limit: Up to 4,096 characters per speech request, according to the documented speech endpoint limit.
  • Audio formats: MP3, Opus, AAC, FLAC, WAV, and PCM.
  • Endpoint: /v1/audio/speech.

The model does not accept images, audio recordings, or video as input. It also does not produce text, images, video, embeddings, or structured JSON as its primary output. The response is audio content, so an application should be prepared to save, stream, play, or further process an audio response rather than handle it like a text completion.

The 4,096-character limit applies to an individual speech request. Longer documents need to be divided into smaller sections before synthesis. Splitting text can also make it easier to manage chapters, paragraphs, pauses, files, and retries, although the supplied research does not specify how applications should preserve continuity between separately generated segments.

Voices and playback controls

The speech endpoint provides preset voices. The documented choices include alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, and cedar. Voice availability can depend on the selected model. OpenAI’s current text-to-speech guide recommends marin and cedar for the best quality, which is a provider recommendation rather than an independent benchmark result.

Applications can also set playback speed. The documented range is 0.25 to 4.0, with a default of 1.0. This allows the same generated content to be made slower for accessibility or language learning, or faster for certain information-delivery use cases. Speed control changes playback timing; it should not be treated as a substitute for rewriting text or controlling pronunciation.

TTS-1 HD does not support the instructions parameter used by newer instruction-steerable text-to-speech models. Developers therefore should not assume that they can reliably direct the model with arbitrary style instructions such as detailed emotional, character, or delivery prompts through that parameter.

Quality, latency, and cost trade-offs

The main reason to choose TTS-1 HD is its quality-oriented positioning. Compared with TTS-1, it is intended for situations where speech quality takes priority over the lowest possible latency. This makes the model a better fit for prepared narration, recorded announcements, accessibility audio, and other content that can tolerate some delay before playback.

The trade-off is not simply audio quality versus speed. TTS-1 HD is also more expensive than the supplied price for TTS-1, because OpenAI lists TTS-1 HD at $30 per 1 million characters for speech generation. The supplied research does not provide a separate output-token price: speech generation is billed by input characters rather than by output tokens.

For a live interaction in which every fraction of a second matters, a lower-latency speech option may be more suitable. For a voice-over pipeline in which audio can be generated in advance, TTS-1 HD’s quality-focused positioning can justify the additional cost or waiting time. The right choice depends on whether the listener values immediate response or polished generated speech more strongly.

Pricing and access

OpenAI lists TTS-1 HD at $30 per 1 million characters. This is character-based speech-generation pricing, not a monthly subscription and not a token price. The practical cost depends on the amount of text sent for synthesis, so applications should account for repeated requests, retries, and any text preprocessing that increases the character count.

The model is accessed through the OpenAI Audio API speech endpoint by specifying tts-1-hd, the text to convert, and a supported voice. Access is subject to OpenAI API availability and account conditions. The supplied research does not specify a free usage allowance, a separate monthly plan, or a maximum number of requests for this model.

Technical capabilities and limitations

TTS-1 HD’s capabilities are focused rather than broad. It supports speech generation and audio streaming, but it is not a tool-using or reasoning model. The research identifies no function-calling support, web search, batch API support, caching, or fine-tuning support for this model.

CapabilityTTS-1 HD
Text inputSupported
Audio outputSupported
Image, audio, or video inputNot supported
Text outputNot the primary output
Reasoning and codingNot supported as model purposes
Tool or function useNot supported
StreamingAudio streaming supported
Maximum input per request4,096 characters
Fine-tuningNot supported

Streaming needs a specific qualification. The speech endpoint supports audio streaming, but server-sent-events streaming is not supported for TTS-1 HD. Applications should use the supported audio streaming format rather than treating the response as an ordinary text event stream.

The model also does not support the newer instruction-based voice controls described for some other text-to-speech options. That limitation matters if an application needs detailed programmatic control over speaking style. TTS-1 HD is better understood as a preset-voice speech generator with playback-speed control than as a fully steerable voice-performance system.

Best use cases for TTS-1 HD

TTS-1 HD is most appropriate when the input is already written and the output needs to be listenable speech. Suitable examples include:

  • Narration and voice-over: Generate spoken versions of scripts, articles, lessons, or prepared presentations.
  • Accessibility: Add read-aloud features for users who prefer or require audio content.
  • Educational content: Produce audio lessons, revision material, or language-learning content from prepared text.
  • Announcements: Create automated informational messages for services, facilities, or applications.
  • Downloadable audio: Export generated speech in a standard format such as MP3, WAV, or FLAC.
  • Quality-focused voice interfaces: Generate responses for an interface when the system can tolerate more latency than a highly time-sensitive conversation.

For longer material, divide the source into requests no larger than the documented 4,096-character input limit. Select the output format based on the destination: compressed formats can be convenient for distribution, while lossless or raw formats may be more appropriate for later audio processing.

When to choose TTS-1 HD

Choose TTS-1 HD when all of the following are true:

  • You already have text and need speech rather than text analysis.
  • Speech quality is more important than the lowest possible latency.
  • A preset voice is sufficient for the application.
  • Your content can be divided into requests of up to 4,096 characters.
  • The $30-per-million-characters price is acceptable for the project.

It may be a strong fit for pre-generated narration, accessibility audio, and announcements because those uses can often generate audio before the listener needs it. The model’s support for several common formats also makes it practical for applications that need to store or distribute generated speech.

When another option may be more appropriate

Choose a lower-latency speech model when the central requirement is immediate conversational response and a quality-oriented model would introduce an unacceptable delay. The sibling TTS-1 is the relevant OpenAI comparison when minimizing latency is more important than selecting the quality-focused TTS-1 HD option.

A newer instruction-steerable text-to-speech option may be more suitable when the application needs detailed control over delivery instructions, because TTS-1 HD does not support the instructions parameter. A model designed for transcription or audio understanding is also required when the task begins with recorded speech rather than text. TTS-1 HD cannot listen to an audio recording and transcribe or analyze it.

Finally, TTS-1 HD is not the right choice for applications that need reasoning, coding, web search, image understanding, or multimodal conversation. Those tasks require a model with the corresponding input, output, and tool capabilities; TTS-1 HD should generally be used as a speech-generation component after another system has produced the text.

Bottom line

TTS-1 HD is a focused OpenAI speech model for converting text into high-quality generated audio. Its key advantages are its quality-oriented position relative to TTS-1, support for multiple audio formats, a broad set of preset voices, adjustable playback speed, and compatibility with the Audio API speech endpoint. Its main constraints are the 4,096-character request limit, character-based pricing of $30 per 1 million characters, lack of image or audio input, lack of reasoning and tool use, and lack of instruction-based voice control.

It is best chosen for prepared or semi-prepared speech content where quality matters more than minimum latency. For highly interactive applications, instruction-heavy voice control, transcription, or general-purpose AI tasks, another model type may be a better fit.


Answers to Frequently Asked Questions

How does TTS-1 HD compare with TTS-1?
TTS-1 HD is positioned as the higher-quality option, while TTS-1 is designed for lower latency. TTS-1 HD may be preferable for prepared narration and other quality-focused applications, whereas TTS-1 may be more suitable for highly interactive experiences where response speed is the priority.
Which voices and audio formats does TTS-1 HD support?
TTS-1 HD supports preset voices including alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, and cedar, although availability may depend on the model. Supported audio formats include MP3, Opus, AAC, FLAC, WAV, and PCM.
What is the maximum text length for a TTS-1 HD request?
TTS-1 HD supports up to 4,096 characters per speech request. Longer documents must be divided into smaller sections before being converted to speech.
What is TTS-1 HD used for?
TTS-1 HD converts written text into generated speech audio through OpenAI’s Audio API. It is best suited for narration, voice-overs, accessibility audio, educational content, announcements, downloadable audio, and quality-focused voice interfaces.
How much does OpenAI TTS-1 HD cost?
OpenAI lists TTS-1 HD at $30 per 1 million characters. Speech generation is billed by input characters rather than output tokens, so repeated requests, retries, and text preprocessing can affect the total cost.


Sources 5
Provider

About OpenAI