What is TTS-1?
TTS-1 is an OpenAI text-to-speech model. It accepts written text and returns spoken audio, rather than generating a text response, image, or video. The model is designed for applications where the time required to begin producing speech matters, such as voice interfaces, web narration, automated announcements, accessibility tools, and other interactive experiences.
OpenAI introduced TTS-1 on November 6, 2023, alongside TTS-1 HD. The two models serve the same broad purpose, but their positioning is different: TTS-1 is optimized for lower latency, while TTS-1 HD is intended to provide higher output quality. This makes TTS-1 a practical choice when an application needs speech to start quickly and does not require the highest-quality option available in the same model family.
TTS-1 is currently listed in OpenAI’s model catalog and speech-generation documentation. It is accessed through the Audio API’s speech endpoint rather than through a general-purpose conversational completion interface.
How TTS-1 works
A developer sends text, selects a voice and output format, and receives an audio file or audio stream. The relevant API operation is the speech-generation endpoint at /v1/audio/speech. TTS-1 can stream audio using chunk transfer encoding, allowing an application to begin playing the result before the entire response has been generated. OpenAI’s documentation states that the server-sent-events stream format is not supported for this model.
The model supports MP3, Opus, AAC, FLAC, WAV, and PCM output. The choice depends on the application: compressed formats such as MP3 can be convenient for delivery, while PCM or WAV may be preferable in workflows that need uncompressed audio. The supplied documentation does not specify a maximum text length, context window, or maximum output duration for TTS-1, so those limits should not be assumed.
Supported input and output types
| Capability | TTS-1 support |
|---|---|
| Text input | Yes |
| Spoken-audio output | Yes |
| Image input | No |
| Audio input | No |
| Video input | No |
| Text output | No; speech is the primary output |
| Image or video output | No |
| Audio streaming | Yes, using chunk transfer encoding |
TTS-1 should therefore be treated as a speech-synthesis component, not as an audio-understanding model. It does not transcribe recordings, interpret spoken commands, answer questions, or reason over documents by itself. An application that needs those capabilities would need a separate model or service before sending text to TTS-1.
Voices and supported languages
The documented built-in voices are alloy, ash, coral, echo, fable, onyx, nova, sage, and shimmer. These voices give developers several preset vocal identities without requiring voice training or fine-tuning.
OpenAI describes the voices as currently optimized for English, while also documenting speech generation across a broad range of languages. That distinction matters for production use: multilingual generation may be available, but the voice quality and pronunciation experience may not be equally strong in every language. Teams should test representative names, technical terms, numbers, abbreviations, and regional pronunciations before selecting a voice for a multilingual product.
TTS-1 pricing
OpenAI lists TTS-1 at $15 per 1 million characters of generated speech. Pricing is based on characters rather than separate input and output token rates. There is no separate token-priced output component in the supplied model information.
For a simple estimate, an application generating 100,000 characters would cost approximately $1.50 at the listed rate, before any other applicable service charges. Actual usage planning should account for repeated requests, retries, introductions or disclaimers added to every response, and any text preprocessing performed by the application.
TTS-1 HD is listed at $30 per 1 million characters, twice the TTS-1 rate. The difference illustrates the principal cost-quality trade-off within this model family: TTS-1 is the lower-cost, lower-latency option, while TTS-1 HD is positioned for higher output quality. These prices are the supplied published model rates and should be checked against OpenAI’s current pricing documentation before deployment.
Speed, quality, and control trade-offs
The clearest provider-supported characteristic of TTS-1 is its low-latency positioning. In practical terms, this is useful when the system should begin speaking promptly after text becomes available. Streaming can reinforce that advantage because the client may start receiving audio while later portions are still being generated.
The trade-off is that TTS-1 is not positioned as OpenAI’s highest-quality speech option. TTS-1 HD is the sibling choice when output quality is more important than the lowest latency or price. OpenAI’s current documentation also recommends the newer GPT-4o mini TTS model for intelligent realtime applications and greater control over accent, emotion, intonation, speed, and tone. That recommendation does not make TTS-1 unusable; it clarifies that TTS-1 is most differentiated by its established speech-endpoint compatibility, predictable character pricing, and low-latency focus.
TTS-1 does not support delivery-style instructions. The current API reference states that the instructions parameter does not work with TTS-1 or TTS-1 HD. Developers should not expect to reliably specify directions such as “speak excitedly,” “use a calm tone,” or “slow down” through that parameter. If expressive control is central to the product, a newer speech model may be more appropriate.
Reasoning, coding, and tool capabilities
TTS-1 is not a general-purpose language model. Its job begins after an application already has text that needs to be spoken. It does not provide general reasoning, code generation, web search, function calling, structured data generation, image generation, or audio understanding as part of the speech-generation task.
It can speak text that was produced by another model or application, including explanations, code-related instructions, customer-service responses, and generated narration. That does not mean TTS-1 understands or evaluates the material. For example, a separate system could analyze a document and produce a summary, then pass the summary to TTS-1 for playback. TTS-1 would synthesize the summary without independently checking its accuracy.
Important limitations
- No documented context or output limits: The supplied research does not provide a context-window size, maximum input length, or maximum output-token limit. Developers should verify operational limits in the current API documentation and test long requests.
- Text-only input: TTS-1 does not accept images, audio, or video as input. It cannot directly describe an image, transcribe a recording, or process a video.
- Speech-only purpose: It is not intended to answer questions, perform reasoning, write code, or generate ordinary text responses.
- Limited delivery control: The
instructionsparameter does not work with TTS-1, which limits explicit control over emotion, tone, accent, intonation, and speaking style. - Voice and language qualification: The voices are currently optimized for English, so multilingual deployments require testing rather than assuming identical quality across languages.
- AI disclosure: Developers should disclose to end users that the voice is AI-generated rather than human.
When to choose TTS-1
Choose TTS-1 when the main requirement is fast, predictable speech generation and the application does not need advanced expressive controls. Good candidates include:
- Realtime-oriented voice interfaces that need responses to begin quickly.
- Website or application narration where low delay is more important than premium voice quality.
- Accessibility features that read text aloud.
- Automated announcements, notifications, and operational messages.
- Large volumes of generated speech where the $15-per-million-character rate is attractive.
- Existing integrations built around the Audio API speech endpoint and its supported output formats.
TTS-1 is less suitable when the product requires highly expressive delivery, fine control of accent or emotion, audio understanding, transcription, or general-purpose reasoning. In those cases, use a model designed for that specific task, and consider GPT-4o mini TTS when the requirement is intelligent realtime speech with more control. TTS-1 HD is the more directly comparable alternative when higher speech quality is preferred over TTS-1’s lower latency and price.
Is TTS-1 the right OpenAI speech model?
TTS-1 remains a sensible option for straightforward text-to-speech pipelines. Its role is easy to define: take text, select one of the documented voices, choose an audio format, and generate speech at a published character-based price. The combination of low-latency positioning, streaming support, and broad format support fits applications where audio needs to be produced reliably and quickly.
It is not the best choice simply because an application contains voice. The correct selection depends on what happens before and after synthesis. If the system only needs to turn prepared text into audio, TTS-1 offers a focused and comparatively economical path. If it must understand speech, control delivery in detail, or produce more refined voice output, a newer or more specialized option may justify its additional cost or different capabilities.

