GPT-4o

GPT-4o Audio

by OpenAI · Retired; API access ended May 7, 2026

GPT-4o Audio was OpenAI’s preview model for spoken conversational applications. It supported text and audio input and output, streaming, function calling, a 128,000-token context window, and 16,384 maximum output tokens. Its audio token pricing was higher than its text pricing, and OpenAI retired the model on May 7, 2026, recommending GPT-Audio 1.5 as its successor.

Text Speech Reasoning Coding
GPT-4o Audio was OpenAI’s audio-capable GPT-4o preview model for voice assistants, spoken customer-service tools, and other applications that needed direct speech understanding and generation. Unlike a text-only model paired with separate transcription and text-to-speech services, it could accept audio and produce audio or text through the Chat Completions API. It is now a historical model rather than an option for new deployments: OpenAI retired the canonical gpt-4o-audio-preview model on May 7, 2026.
Outputs

What GPT-4o Audio can produce

Text Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Tool use Streaming Batch API Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
6/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o
Model type Multimodal
Context window 128K tokens
Maximum output 16K tokens
Knowledge cutoff 2023-10-01
Release date 2024-10-17
Status Retired; API access ended May 7, 2026
Deprecation date 2025-09-15
Shutdown date 2026-05-07
Knowledge cutoff notes

The official model documentation lists October 1, 2023 as the knowledge cutoff. This cutoff applies to the underlying model and is not changed by external tools, retrieval, or application-provided context.

Model notes

The canonical API identifier was gpt-4o-audio-preview. It was a preview model available through Chat Completions and related documented endpoints. The model supported text and audio input/output, streaming, and function calling, but not structured outputs or fine-tuning. Documented snapshots included gpt-4o-audio-preview-2024-10-01, gpt-4o-audio-preview-2024-12-17, and gpt-4o-audio-preview-2025-06-03. OpenAI announced GPT-Audio 1.5 as the replacement. The October 1, 2024 snapshot was retired earlier, on October 10, 2025.

Cost

Model pricing

Input $2.50 per 1M text input tokens; $40 per 1M audio input tokens
Output $10.00 per 1M text output tokens; $80 per 1M audio output tokens
Model guide

GPT-4o Audio: OpenAI’s Retired Voice Conversation Model

GPT-4o Audio was OpenAI’s preview model for applications that needed to understand audio and return spoken or written responses. Available through the Chat Completions API under the gpt-4o-audio-preview identifier, it supported text and audio input, text and audio output, streaming, and function calling. It offered a 128,000-token context window and a 16,384-token maximum output, but it did not support images, video, structured outputs, or fine-tuning. OpenAI retired the canonical model on May 7, 2026 and recommended GPT-Audio 1.5 as its replacement.

What GPT-4o Audio was

GPT-4o Audio was a preview model from OpenAI designed for conversational applications involving speech. Its canonical API identifier was gpt-4o-audio-preview. The model accepted either text or audio as input and could respond with either text or audio, allowing developers to build voice interactions without treating transcription, language understanding, and speech generation as entirely separate stages.

OpenAI introduced the model for the Chat Completions API on October 17, 2024. The model used the same underlying model as OpenAI’s Realtime API audio experience, but the documented GPT-4o Audio offering was exposed as a preview model through Chat Completions and related endpoints.

GPT-4o Audio is no longer available for new API use. OpenAI announced that the canonical model would be removed on May 7, 2026, and listed GPT-Audio 1.5 as the recommended replacement. That retirement matters more than the model’s historical capabilities when evaluating it for a current production system.

Capabilities and technical limits

The model’s main distinction was its combination of audio understanding and audio generation. It could interpret spoken input and return a spoken conversational answer, while also supporting ordinary text input and output. It did not accept images or video, and it did not generate images or video.

SpecificationDocumented behavior
API identifiergpt-4o-audio-preview
Context window128,000 tokens
Maximum output16,384 tokens
Input modalitiesText and audio
Output modalitiesText and audio
StreamingSupported
Function callingSupported
Structured outputsNot supported
Fine-tuningNot supported
Knowledge cutoffOctober 1, 2023

The context window is the amount of tokenized conversation and other input the model could consider at once. Its 128,000-token capacity was substantial for long conversations, although audio usage was billed using audio tokens rather than the same text-token rate. The 16,384-token output limit describes the maximum model response, not a guarantee that an audio conversation would produce a response of that length.

Function calling allowed the model to request an application-defined operation, such as looking up an account or retrieving information from a business system. However, function calling should not be confused with structured outputs: the supplied documentation identifies structured outputs as unsupported, so workflows requiring strict JSON Schema enforcement were not a suitable fit.

Audio and text workflows

GPT-4o Audio was intended for applications where speech was part of the interaction itself. A voice assistant could receive a spoken question and answer aloud. A customer-service application could accept a caller’s speech, use function calling to invoke an application operation, and return a spoken response. A language-practice tool could use the model’s audio input and output for conversational exercises.

The model could also return text rather than audio. That was useful when an application needed a transcript-like answer, text for a user interface, or a text response that another system would process. Conversely, audio input and audio output were useful when the goal was a natural spoken exchange rather than a conventional chat window.

Its audio output was designed for spoken conversational responses. The supplied specifications do not describe it as a music-generation model, and music generation was not an appropriate use case.

Historical API pricing

GPT-4o Audio used separate rates for text and audio tokens. Before retirement, the documented prices were:

  • Text input: $2.50 per 1 million tokens.
  • Text output: $10 per 1 million tokens.
  • Audio input: $40 per 1 million tokens.
  • Audio output: $80 per 1 million tokens.

Audio processing was therefore substantially more expensive per million tokens than text processing. A system that routed every interaction through audio could cost more than one that used audio selectively or returned text when spoken output was unnecessary. These are historical API prices only; they should not be treated as current purchasing options because the model was retired.

Strengths and trade-offs

The model’s strongest practical feature was modality integration. Developers could use one model for spoken input, language understanding, tool requests, and spoken or written responses. That could simplify a voice experience compared with assembling independent speech-recognition, language-model, and text-to-speech components.

It also combined a large 128,000-token context window with streaming and function calling. Those features were useful for interactive conversations where responses needed to arrive incrementally and where the assistant had to interact with application systems.

The trade-offs were equally important. Audio tokens were priced much higher than text tokens, the model did not support structured outputs or fine-tuning, and its preview status meant that the model was not a stable long-term choice. It also lacked image and video capabilities. Applications requiring strict machine-readable schemas, visual understanding, custom fine-tuning, or lower-cost text processing needed a different model type.

The supplied evaluation data assigns GPT-4o Audio editorial scores of 7 for reasoning, 7 for coding, 6 for speed, and 4 for cost. These are comparative editorial assessments, not OpenAI-published benchmark results. They suggest a model that was considered reasonably capable for reasoning and coding tasks, but whose audio pricing reduced its cost attractiveness.

Best use cases

  • Voice assistants: spoken questions and spoken answers in a conversational interface.
  • Spoken customer service: audio interactions combined with application functions and business data.
  • Audio-enabled agents: assistants that need to understand speech and respond without a separate speech pipeline.
  • Language practice: interactive spoken exercises with text or audio responses.
  • Hybrid interfaces: applications that accept audio but sometimes return text, or accept text and respond aloud.

These use cases describe where the model fit before retirement. They do not imply that GPT-4o Audio should be selected for a new deployment today.

When another option is more appropriate

For a new voice application, the most important alternative is the provider’s recommended successor, GPT-Audio 1.5. The supplied research does not provide a full specification or price comparison for that successor, so detailed performance or cost claims would be unwarranted. It is nevertheless the relevant direction for replacing GPT-4o Audio because OpenAI named it as the replacement after retirement.

A text-first model may be more appropriate when the application does not need spoken output, particularly because GPT-4o Audio’s audio token rates were high. A dedicated text-to-speech or transcription pipeline may also be preferable when an application needs independently replaceable speech components, although the supplied research does not establish comparative performance for those alternatives.

GPT-4o Audio was not the right choice for image or video workflows, strict JSON Schema output, model fine-tuning, or music generation. Those requirements fall outside its documented capabilities and should be treated as selection blockers rather than minor limitations.

Availability and retirement history

OpenAI documented several dated snapshots, including gpt-4o-audio-preview-2024-10-01, gpt-4o-audio-preview-2024-12-17, and gpt-4o-audio-preview-2025-06-03. The October 1, 2024 snapshot had an earlier retirement date of October 10, 2025. That snapshot-level shutdown should be distinguished from the later retirement of the canonical gpt-4o-audio-preview alias on May 7, 2026.

In summary, GPT-4o Audio was a specialized preview model for direct audio conversation: capable of accepting and generating speech, supporting streaming and function calling, and handling long contexts. Its historical value is clearest for understanding OpenAI’s audio-model progression, while current implementations should follow the provider’s stated successor rather than attempt to start a new integration with this retired model.


Answers to Frequently Asked Questions

What should developers use instead of GPT-4o Audio?
For a new voice application, OpenAI identified GPT-Audio 1.5 as the recommended successor. A text-first model or a separate transcription and text-to-speech pipeline may be more suitable when spoken interaction is not required.
How much did GPT-4o Audio API usage cost?
Its historical pricing was $2.50 per million text input tokens, $10 per million text output tokens, $40 per million audio input tokens, and $80 per million audio output tokens. These prices are no longer current purchasing options because the model was retired.
What were GPT-4o Audio’s main capabilities and limitations?
GPT-4o Audio supported audio and text inputs and outputs, streaming, function calling, and a 128,000-token context window. It did not support images, video, structured outputs, fine-tuning, or music generation.
What was GPT-4o Audio?
GPT-4o Audio was OpenAI’s preview model for conversational applications involving speech. Its API identifier was gpt-4o-audio-preview, and it could accept text or audio input and return text or audio output.
Is GPT-4o Audio still available for new API projects?
No. GPT-4o Audio is no longer available for new API use. OpenAI announced that the canonical gpt-4o-audio-preview model would be removed on May 7, 2026, with GPT-Audio 1.5 recommended as its replacement.


Sources 4
Provider

About OpenAI