What GPT-4o Audio was
GPT-4o Audio was a preview model from OpenAI designed for conversational applications involving speech. Its canonical API identifier was gpt-4o-audio-preview. The model accepted either text or audio as input and could respond with either text or audio, allowing developers to build voice interactions without treating transcription, language understanding, and speech generation as entirely separate stages.
OpenAI introduced the model for the Chat Completions API on October 17, 2024. The model used the same underlying model as OpenAI’s Realtime API audio experience, but the documented GPT-4o Audio offering was exposed as a preview model through Chat Completions and related endpoints.
GPT-4o Audio is no longer available for new API use. OpenAI announced that the canonical model would be removed on May 7, 2026, and listed GPT-Audio 1.5 as the recommended replacement. That retirement matters more than the model’s historical capabilities when evaluating it for a current production system.
Capabilities and technical limits
The model’s main distinction was its combination of audio understanding and audio generation. It could interpret spoken input and return a spoken conversational answer, while also supporting ordinary text input and output. It did not accept images or video, and it did not generate images or video.
| Specification | Documented behavior |
|---|---|
| API identifier | gpt-4o-audio-preview |
| Context window | 128,000 tokens |
| Maximum output | 16,384 tokens |
| Input modalities | Text and audio |
| Output modalities | Text and audio |
| Streaming | Supported |
| Function calling | Supported |
| Structured outputs | Not supported |
| Fine-tuning | Not supported |
| Knowledge cutoff | October 1, 2023 |
The context window is the amount of tokenized conversation and other input the model could consider at once. Its 128,000-token capacity was substantial for long conversations, although audio usage was billed using audio tokens rather than the same text-token rate. The 16,384-token output limit describes the maximum model response, not a guarantee that an audio conversation would produce a response of that length.
Function calling allowed the model to request an application-defined operation, such as looking up an account or retrieving information from a business system. However, function calling should not be confused with structured outputs: the supplied documentation identifies structured outputs as unsupported, so workflows requiring strict JSON Schema enforcement were not a suitable fit.
Audio and text workflows
GPT-4o Audio was intended for applications where speech was part of the interaction itself. A voice assistant could receive a spoken question and answer aloud. A customer-service application could accept a caller’s speech, use function calling to invoke an application operation, and return a spoken response. A language-practice tool could use the model’s audio input and output for conversational exercises.
The model could also return text rather than audio. That was useful when an application needed a transcript-like answer, text for a user interface, or a text response that another system would process. Conversely, audio input and audio output were useful when the goal was a natural spoken exchange rather than a conventional chat window.
Its audio output was designed for spoken conversational responses. The supplied specifications do not describe it as a music-generation model, and music generation was not an appropriate use case.
Historical API pricing
GPT-4o Audio used separate rates for text and audio tokens. Before retirement, the documented prices were:
- Text input: $2.50 per 1 million tokens.
- Text output: $10 per 1 million tokens.
- Audio input: $40 per 1 million tokens.
- Audio output: $80 per 1 million tokens.
Audio processing was therefore substantially more expensive per million tokens than text processing. A system that routed every interaction through audio could cost more than one that used audio selectively or returned text when spoken output was unnecessary. These are historical API prices only; they should not be treated as current purchasing options because the model was retired.
Strengths and trade-offs
The model’s strongest practical feature was modality integration. Developers could use one model for spoken input, language understanding, tool requests, and spoken or written responses. That could simplify a voice experience compared with assembling independent speech-recognition, language-model, and text-to-speech components.
It also combined a large 128,000-token context window with streaming and function calling. Those features were useful for interactive conversations where responses needed to arrive incrementally and where the assistant had to interact with application systems.
The trade-offs were equally important. Audio tokens were priced much higher than text tokens, the model did not support structured outputs or fine-tuning, and its preview status meant that the model was not a stable long-term choice. It also lacked image and video capabilities. Applications requiring strict machine-readable schemas, visual understanding, custom fine-tuning, or lower-cost text processing needed a different model type.
The supplied evaluation data assigns GPT-4o Audio editorial scores of 7 for reasoning, 7 for coding, 6 for speed, and 4 for cost. These are comparative editorial assessments, not OpenAI-published benchmark results. They suggest a model that was considered reasonably capable for reasoning and coding tasks, but whose audio pricing reduced its cost attractiveness.
Best use cases
- Voice assistants: spoken questions and spoken answers in a conversational interface.
- Spoken customer service: audio interactions combined with application functions and business data.
- Audio-enabled agents: assistants that need to understand speech and respond without a separate speech pipeline.
- Language practice: interactive spoken exercises with text or audio responses.
- Hybrid interfaces: applications that accept audio but sometimes return text, or accept text and respond aloud.
These use cases describe where the model fit before retirement. They do not imply that GPT-4o Audio should be selected for a new deployment today.
When another option is more appropriate
For a new voice application, the most important alternative is the provider’s recommended successor, GPT-Audio 1.5. The supplied research does not provide a full specification or price comparison for that successor, so detailed performance or cost claims would be unwarranted. It is nevertheless the relevant direction for replacing GPT-4o Audio because OpenAI named it as the replacement after retirement.
A text-first model may be more appropriate when the application does not need spoken output, particularly because GPT-4o Audio’s audio token rates were high. A dedicated text-to-speech or transcription pipeline may also be preferable when an application needs independently replaceable speech components, although the supplied research does not establish comparative performance for those alternatives.
GPT-4o Audio was not the right choice for image or video workflows, strict JSON Schema output, model fine-tuning, or music generation. Those requirements fall outside its documented capabilities and should be treated as selection blockers rather than minor limitations.
Availability and retirement history
OpenAI documented several dated snapshots, including gpt-4o-audio-preview-2024-10-01, gpt-4o-audio-preview-2024-12-17, and gpt-4o-audio-preview-2025-06-03. The October 1, 2024 snapshot had an earlier retirement date of October 10, 2025. That snapshot-level shutdown should be distinguished from the later retirement of the canonical gpt-4o-audio-preview alias on May 7, 2026.
In summary, GPT-4o Audio was a specialized preview model for direct audio conversation: capable of accepting and generating speech, supporting streaming and function calling, and handling long contexts. Its historical value is clearest for understanding OpenAI’s audio-model progression, while current implementations should follow the provider’s stated successor rather than attempt to start a new integration with this retired model.

