TTS-1

TTS-1

by OpenAI · Current and accessible; optimized for low-latency text-to-speech

OpenAI TTS-1 converts text into spoken audio with a focus on low latency. It supports nine built-in voices, MP3, Opus, AAC, FLAC, WAV and PCM output, chunked streaming, and pricing of $15 per 1 million generated characters. It is best for straightforward, fast speech synthesis rather than audio understanding or highly expressive voice control.

Speech Reasoning Coding
TTS-1 is OpenAI’s speech-generation model for applications that need spoken audio quickly. It turns text into natural-sounding speech through the Audio API, making it suitable for voice interfaces, narration, accessibility features, announcements, and other realtime-oriented experiences. Its main trade-off is straightforward: TTS-1 prioritizes latency and cost over the highest available voice quality and advanced control over delivery style.
Outputs

What TTS-1 can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family TTS-1
Model type Other
Release date 2023-11-06
Status Current and accessible; optimized for low-latency text-to-speech
Knowledge cutoff notes

OpenAI does not publish a knowledge cutoff for TTS-1. As a text-to-speech model, its primary function is speech generation rather than knowledge retrieval or general text completion.

Model notes

TTS-1 is a speech-generation model rather than a general-purpose language model. It accepts text and returns spoken audio through the Audio API speech endpoint. OpenAI documents support for MP3, Opus, AAC, FLAC, WAV, and PCM output. The model supports the voices alloy, ash, coral, echo, fable, onyx, nova, sage, and shimmer. The instructions parameter does not work with TTS-1 or TTS-1 HD. Pricing is based on generated characters rather than separate input and output token rates. OpenAI's current documentation recommends GPT-4o mini TTS for newer intelligent realtime applications, but TTS-1 remains listed as an accessible model.

Cost

Model pricing

Input $15.00 per 1 million characters
Model guide

TTS-1: OpenAI’s Low-Latency Text-to-Speech Model, Pricing and Capabilities

TTS-1 is OpenAI’s text-to-speech model for converting written text into spoken audio with low response latency. It is available through the Audio API speech endpoint, supports several built-in voices and audio formats, can stream audio, and is priced at $15 per 1 million generated characters.

What is TTS-1?

TTS-1 is an OpenAI text-to-speech model. It accepts written text and returns spoken audio, rather than generating a text response, image, or video. The model is designed for applications where the time required to begin producing speech matters, such as voice interfaces, web narration, automated announcements, accessibility tools, and other interactive experiences.

OpenAI introduced TTS-1 on November 6, 2023, alongside TTS-1 HD. The two models serve the same broad purpose, but their positioning is different: TTS-1 is optimized for lower latency, while TTS-1 HD is intended to provide higher output quality. This makes TTS-1 a practical choice when an application needs speech to start quickly and does not require the highest-quality option available in the same model family.

TTS-1 is currently listed in OpenAI’s model catalog and speech-generation documentation. It is accessed through the Audio API’s speech endpoint rather than through a general-purpose conversational completion interface.

How TTS-1 works

A developer sends text, selects a voice and output format, and receives an audio file or audio stream. The relevant API operation is the speech-generation endpoint at /v1/audio/speech. TTS-1 can stream audio using chunk transfer encoding, allowing an application to begin playing the result before the entire response has been generated. OpenAI’s documentation states that the server-sent-events stream format is not supported for this model.

The model supports MP3, Opus, AAC, FLAC, WAV, and PCM output. The choice depends on the application: compressed formats such as MP3 can be convenient for delivery, while PCM or WAV may be preferable in workflows that need uncompressed audio. The supplied documentation does not specify a maximum text length, context window, or maximum output duration for TTS-1, so those limits should not be assumed.

Supported input and output types

CapabilityTTS-1 support
Text inputYes
Spoken-audio outputYes
Image inputNo
Audio inputNo
Video inputNo
Text outputNo; speech is the primary output
Image or video outputNo
Audio streamingYes, using chunk transfer encoding

TTS-1 should therefore be treated as a speech-synthesis component, not as an audio-understanding model. It does not transcribe recordings, interpret spoken commands, answer questions, or reason over documents by itself. An application that needs those capabilities would need a separate model or service before sending text to TTS-1.

Voices and supported languages

The documented built-in voices are alloy, ash, coral, echo, fable, onyx, nova, sage, and shimmer. These voices give developers several preset vocal identities without requiring voice training or fine-tuning.

OpenAI describes the voices as currently optimized for English, while also documenting speech generation across a broad range of languages. That distinction matters for production use: multilingual generation may be available, but the voice quality and pronunciation experience may not be equally strong in every language. Teams should test representative names, technical terms, numbers, abbreviations, and regional pronunciations before selecting a voice for a multilingual product.

TTS-1 pricing

OpenAI lists TTS-1 at $15 per 1 million characters of generated speech. Pricing is based on characters rather than separate input and output token rates. There is no separate token-priced output component in the supplied model information.

For a simple estimate, an application generating 100,000 characters would cost approximately $1.50 at the listed rate, before any other applicable service charges. Actual usage planning should account for repeated requests, retries, introductions or disclaimers added to every response, and any text preprocessing performed by the application.

TTS-1 HD is listed at $30 per 1 million characters, twice the TTS-1 rate. The difference illustrates the principal cost-quality trade-off within this model family: TTS-1 is the lower-cost, lower-latency option, while TTS-1 HD is positioned for higher output quality. These prices are the supplied published model rates and should be checked against OpenAI’s current pricing documentation before deployment.

Speed, quality, and control trade-offs

The clearest provider-supported characteristic of TTS-1 is its low-latency positioning. In practical terms, this is useful when the system should begin speaking promptly after text becomes available. Streaming can reinforce that advantage because the client may start receiving audio while later portions are still being generated.

The trade-off is that TTS-1 is not positioned as OpenAI’s highest-quality speech option. TTS-1 HD is the sibling choice when output quality is more important than the lowest latency or price. OpenAI’s current documentation also recommends the newer GPT-4o mini TTS model for intelligent realtime applications and greater control over accent, emotion, intonation, speed, and tone. That recommendation does not make TTS-1 unusable; it clarifies that TTS-1 is most differentiated by its established speech-endpoint compatibility, predictable character pricing, and low-latency focus.

TTS-1 does not support delivery-style instructions. The current API reference states that the instructions parameter does not work with TTS-1 or TTS-1 HD. Developers should not expect to reliably specify directions such as “speak excitedly,” “use a calm tone,” or “slow down” through that parameter. If expressive control is central to the product, a newer speech model may be more appropriate.

Reasoning, coding, and tool capabilities

TTS-1 is not a general-purpose language model. Its job begins after an application already has text that needs to be spoken. It does not provide general reasoning, code generation, web search, function calling, structured data generation, image generation, or audio understanding as part of the speech-generation task.

It can speak text that was produced by another model or application, including explanations, code-related instructions, customer-service responses, and generated narration. That does not mean TTS-1 understands or evaluates the material. For example, a separate system could analyze a document and produce a summary, then pass the summary to TTS-1 for playback. TTS-1 would synthesize the summary without independently checking its accuracy.

Important limitations

  • No documented context or output limits: The supplied research does not provide a context-window size, maximum input length, or maximum output-token limit. Developers should verify operational limits in the current API documentation and test long requests.
  • Text-only input: TTS-1 does not accept images, audio, or video as input. It cannot directly describe an image, transcribe a recording, or process a video.
  • Speech-only purpose: It is not intended to answer questions, perform reasoning, write code, or generate ordinary text responses.
  • Limited delivery control: The instructions parameter does not work with TTS-1, which limits explicit control over emotion, tone, accent, intonation, and speaking style.
  • Voice and language qualification: The voices are currently optimized for English, so multilingual deployments require testing rather than assuming identical quality across languages.
  • AI disclosure: Developers should disclose to end users that the voice is AI-generated rather than human.

When to choose TTS-1

Choose TTS-1 when the main requirement is fast, predictable speech generation and the application does not need advanced expressive controls. Good candidates include:

  • Realtime-oriented voice interfaces that need responses to begin quickly.
  • Website or application narration where low delay is more important than premium voice quality.
  • Accessibility features that read text aloud.
  • Automated announcements, notifications, and operational messages.
  • Large volumes of generated speech where the $15-per-million-character rate is attractive.
  • Existing integrations built around the Audio API speech endpoint and its supported output formats.

TTS-1 is less suitable when the product requires highly expressive delivery, fine control of accent or emotion, audio understanding, transcription, or general-purpose reasoning. In those cases, use a model designed for that specific task, and consider GPT-4o mini TTS when the requirement is intelligent realtime speech with more control. TTS-1 HD is the more directly comparable alternative when higher speech quality is preferred over TTS-1’s lower latency and price.

Is TTS-1 the right OpenAI speech model?

TTS-1 remains a sensible option for straightforward text-to-speech pipelines. Its role is easy to define: take text, select one of the documented voices, choose an audio format, and generate speech at a published character-based price. The combination of low-latency positioning, streaming support, and broad format support fits applications where audio needs to be produced reliably and quickly.

It is not the best choice simply because an application contains voice. The correct selection depends on what happens before and after synthesis. If the system only needs to turn prepared text into audio, TTS-1 offers a focused and comparatively economical path. If it must understand speech, control delivery in detail, or produce more refined voice output, a newer or more specialized option may justify its additional cost or different capabilities.


Answers to Frequently Asked Questions

How does TTS-1 compare with TTS-1 HD?
TTS-1 is optimized for lower latency and costs $15 per 1 million characters, while TTS-1 HD is positioned for higher speech quality and costs $30 per 1 million characters. TTS-1 is generally preferable when fast, economical speech generation is more important than premium output quality.
Does TTS-1 support real-time audio streaming?
Yes. TTS-1 can stream audio using chunk transfer encoding, allowing an application to start playing speech before the entire response has been generated. The server-sent-events stream format is not supported.
How much does OpenAI TTS-1 cost?
TTS-1 costs $15 per 1 million generated characters. For example, generating 100,000 characters would cost approximately $1.50 before any other applicable service charges.
Which voices and audio formats does TTS-1 support?
TTS-1 supports the voices alloy, ash, coral, echo, fable, onyx, nova, sage, and shimmer. It can generate audio in MP3, Opus, AAC, FLAC, WAV, and PCM formats.
What is OpenAI TTS-1 used for?
TTS-1 is an OpenAI text-to-speech model that converts written text into spoken audio. It is designed for low-latency applications such as voice interfaces, website narration, accessibility tools, automated announcements, notifications, and other interactive experiences.


Sources 4
Provider

About OpenAI