What is TTS-1 HD?
TTS-1 HD is OpenAI’s quality-focused text-to-speech model. Text-to-speech, often abbreviated as TTS, converts written words into spoken audio. TTS-1 HD performs this conversion through OpenAI’s Audio API speech endpoint using the model identifier tts-1-hd.
The model is positioned as the higher-quality counterpart to TTS-1. The distinction is mainly about the trade-off between audio quality and latency: TTS-1 is designed for lower-latency speech generation, while TTS-1 HD is intended for cases where a more quality-oriented result is worth accepting potentially higher latency or cost.
TTS-1 HD is not a general-purpose language model. Its job is to generate speech from supplied text. It is not designed for reasoning, coding, transcription, image analysis, web search, or general conversation by itself.
Where TTS-1 HD fits in OpenAI’s catalog
TTS-1 HD is part of OpenAI’s Audio API model lineup and is currently documented as an available quality-oriented speech model. OpenAI introduced TTS-1 HD on November 6, 2023, alongside TTS-1. The current model catalog uses tts-1-hd as its canonical API identifier.
OpenAI’s newer text-to-speech documentation recommends GPT-4o mini TTS for intelligent real-time applications, but it continues to document TTS-1 HD as a speech-generation option. In practical terms, TTS-1 HD remains relevant when an application already has text and needs downloadable or streamed speech, particularly when quality is more important than the minimum possible delay.
Inputs, outputs, and supported modalities
TTS-1 HD has a narrow, clearly defined modality profile:
- Input: Text.
- Output: Generated speech audio.
- Input limit: Up to 4,096 characters per speech request, according to the documented speech endpoint limit.
- Audio formats: MP3, Opus, AAC, FLAC, WAV, and PCM.
- Endpoint:
/v1/audio/speech.
The model does not accept images, audio recordings, or video as input. It also does not produce text, images, video, embeddings, or structured JSON as its primary output. The response is audio content, so an application should be prepared to save, stream, play, or further process an audio response rather than handle it like a text completion.
The 4,096-character limit applies to an individual speech request. Longer documents need to be divided into smaller sections before synthesis. Splitting text can also make it easier to manage chapters, paragraphs, pauses, files, and retries, although the supplied research does not specify how applications should preserve continuity between separately generated segments.
Voices and playback controls
The speech endpoint provides preset voices. The documented choices include alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, and cedar. Voice availability can depend on the selected model. OpenAI’s current text-to-speech guide recommends marin and cedar for the best quality, which is a provider recommendation rather than an independent benchmark result.
Applications can also set playback speed. The documented range is 0.25 to 4.0, with a default of 1.0. This allows the same generated content to be made slower for accessibility or language learning, or faster for certain information-delivery use cases. Speed control changes playback timing; it should not be treated as a substitute for rewriting text or controlling pronunciation.
TTS-1 HD does not support the instructions parameter used by newer instruction-steerable text-to-speech models. Developers therefore should not assume that they can reliably direct the model with arbitrary style instructions such as detailed emotional, character, or delivery prompts through that parameter.
Quality, latency, and cost trade-offs
The main reason to choose TTS-1 HD is its quality-oriented positioning. Compared with TTS-1, it is intended for situations where speech quality takes priority over the lowest possible latency. This makes the model a better fit for prepared narration, recorded announcements, accessibility audio, and other content that can tolerate some delay before playback.
The trade-off is not simply audio quality versus speed. TTS-1 HD is also more expensive than the supplied price for TTS-1, because OpenAI lists TTS-1 HD at $30 per 1 million characters for speech generation. The supplied research does not provide a separate output-token price: speech generation is billed by input characters rather than by output tokens.
For a live interaction in which every fraction of a second matters, a lower-latency speech option may be more suitable. For a voice-over pipeline in which audio can be generated in advance, TTS-1 HD’s quality-focused positioning can justify the additional cost or waiting time. The right choice depends on whether the listener values immediate response or polished generated speech more strongly.
Pricing and access
OpenAI lists TTS-1 HD at $30 per 1 million characters. This is character-based speech-generation pricing, not a monthly subscription and not a token price. The practical cost depends on the amount of text sent for synthesis, so applications should account for repeated requests, retries, and any text preprocessing that increases the character count.
The model is accessed through the OpenAI Audio API speech endpoint by specifying tts-1-hd, the text to convert, and a supported voice. Access is subject to OpenAI API availability and account conditions. The supplied research does not specify a free usage allowance, a separate monthly plan, or a maximum number of requests for this model.
Technical capabilities and limitations
TTS-1 HD’s capabilities are focused rather than broad. It supports speech generation and audio streaming, but it is not a tool-using or reasoning model. The research identifies no function-calling support, web search, batch API support, caching, or fine-tuning support for this model.
| Capability | TTS-1 HD |
|---|---|
| Text input | Supported |
| Audio output | Supported |
| Image, audio, or video input | Not supported |
| Text output | Not the primary output |
| Reasoning and coding | Not supported as model purposes |
| Tool or function use | Not supported |
| Streaming | Audio streaming supported |
| Maximum input per request | 4,096 characters |
| Fine-tuning | Not supported |
Streaming needs a specific qualification. The speech endpoint supports audio streaming, but server-sent-events streaming is not supported for TTS-1 HD. Applications should use the supported audio streaming format rather than treating the response as an ordinary text event stream.
The model also does not support the newer instruction-based voice controls described for some other text-to-speech options. That limitation matters if an application needs detailed programmatic control over speaking style. TTS-1 HD is better understood as a preset-voice speech generator with playback-speed control than as a fully steerable voice-performance system.
Best use cases for TTS-1 HD
TTS-1 HD is most appropriate when the input is already written and the output needs to be listenable speech. Suitable examples include:
- Narration and voice-over: Generate spoken versions of scripts, articles, lessons, or prepared presentations.
- Accessibility: Add read-aloud features for users who prefer or require audio content.
- Educational content: Produce audio lessons, revision material, or language-learning content from prepared text.
- Announcements: Create automated informational messages for services, facilities, or applications.
- Downloadable audio: Export generated speech in a standard format such as MP3, WAV, or FLAC.
- Quality-focused voice interfaces: Generate responses for an interface when the system can tolerate more latency than a highly time-sensitive conversation.
For longer material, divide the source into requests no larger than the documented 4,096-character input limit. Select the output format based on the destination: compressed formats can be convenient for distribution, while lossless or raw formats may be more appropriate for later audio processing.
When to choose TTS-1 HD
Choose TTS-1 HD when all of the following are true:
- You already have text and need speech rather than text analysis.
- Speech quality is more important than the lowest possible latency.
- A preset voice is sufficient for the application.
- Your content can be divided into requests of up to 4,096 characters.
- The $30-per-million-characters price is acceptable for the project.
It may be a strong fit for pre-generated narration, accessibility audio, and announcements because those uses can often generate audio before the listener needs it. The model’s support for several common formats also makes it practical for applications that need to store or distribute generated speech.
When another option may be more appropriate
Choose a lower-latency speech model when the central requirement is immediate conversational response and a quality-oriented model would introduce an unacceptable delay. The sibling TTS-1 is the relevant OpenAI comparison when minimizing latency is more important than selecting the quality-focused TTS-1 HD option.
A newer instruction-steerable text-to-speech option may be more suitable when the application needs detailed control over delivery instructions, because TTS-1 HD does not support the instructions parameter. A model designed for transcription or audio understanding is also required when the task begins with recorded speech rather than text. TTS-1 HD cannot listen to an audio recording and transcribe or analyze it.
Finally, TTS-1 HD is not the right choice for applications that need reasoning, coding, web search, image understanding, or multimodal conversation. Those tasks require a model with the corresponding input, output, and tool capabilities; TTS-1 HD should generally be used as a speech-generation component after another system has produced the text.
Bottom line
TTS-1 HD is a focused OpenAI speech model for converting text into high-quality generated audio. Its key advantages are its quality-oriented position relative to TTS-1, support for multiple audio formats, a broad set of preset voices, adjustable playback speed, and compatibility with the Audio API speech endpoint. Its main constraints are the 4,096-character request limit, character-based pricing of $30 per 1 million characters, lack of image or audio input, lack of reasoning and tool use, and lack of instruction-based voice control.
It is best chosen for prepared or semi-prepared speech content where quality matters more than minimum latency. For highly interactive applications, instruction-heavy voice control, transcription, or general-purpose AI tasks, another model type may be a better fit.

