What is GPT-Audio?
GPT-Audio is an OpenAI multimodal model built for audio conversations in the Chat Completions API. In practical terms, an application can send the model written text, recorded or streamed audio, or a combination of text and audio, and receive either text or generated audio in response. This makes it suitable for conversational experiences where users speak naturally and the application answers aloud.
The model is not a general-purpose image-and-video model. Its documented modalities are text and audio: it can accept text and audio inputs and produce text and audio outputs. OpenAI describes it as its first generally available audio model, with support for direct audio understanding and generation through Chat Completions.
The canonical model identifier is gpt-audio. OpenAI also documents the snapshot identifier gpt-audio-2025-08-28. The model was released on August 28, 2025, according to the supplied model data.
Where GPT-Audio fits in OpenAI’s lineup
GPT-Audio occupies a specialized position rather than serving as a universal model for every media type. Its purpose is direct audio interaction inside the Chat Completions workflow. That distinguishes it from text-only models, which require a separate speech-recognition step before the language model can process a spoken request, and from models intended for advanced real-time voice-agent workflows.
That position is now transitional. OpenAI lists GPT-Audio as deprecated, with a documented deprecation date of July 20, 2026, and a shutdown date of January 20, 2027. OpenAI recommends migrating to GPT-Audio-1.5 before the shutdown. GPT-Audio can therefore still be relevant when maintaining an existing integration or testing behavior tied specifically to this model, but it is not the safest choice for a new application expected to operate beyond the shutdown date.
Core capabilities and supported modalities
GPT-Audio’s defining capability is native handling of both audio and text within a conversational model request. Audio input can represent a spoken user request or other audio content the application needs the model to interpret. Audio output allows the response to be delivered as speech rather than only as text.
| Capability | GPT-Audio support |
|---|---|
| Text input | Supported |
| Audio input | Supported |
| Text output | Supported |
| Audio or speech output | Supported |
| Image input | Not supported |
| Video input | Not supported |
| Streaming | Supported |
| Function calling | Supported |
| Fine-tuning | Not supported |
| Structured output | Not supported |
Streaming is important for voice experiences because it can allow an application to begin handling or delivering a response before the entire exchange is complete. Function calling allows the model to request application-defined operations, such as looking up information or triggering a workflow. The model itself does not automatically perform arbitrary actions; the surrounding application must implement and authorize the available functions.
Context window and output limit
GPT-Audio has a documented context length of 128,000 tokens. A token is a unit of text or model-processing data, so the context window represents the amount of conversation and other supported content the model can consider within a request. A larger context can help when an audio conversation includes substantial accompanying text or a long prior exchange, although the usable capacity depends on the complete request and the way audio is represented.
The maximum output limit is 16,384 tokens. This is a ceiling on the generated response rather than a promise that every request will produce an output of that size. In a spoken interface, practical response length should usually be controlled by the application so that answers remain timely and easy to follow.
The model’s listed knowledge cutoff is October 1, 2023. Sending audio, adding external context, or connecting application-level retrieval does not change that underlying cutoff. If an application needs current information, it must supply that information through its own tools or external systems; GPT-Audio is not documented as supporting web search directly.
GPT-Audio pricing
GPT-Audio uses separate token rates for text and audio, with different prices for input and output. The supplied OpenAI pricing information is:
| Token type | Input price per 1 million tokens | Output price per 1 million tokens |
|---|---|---|
| Text | $2.50 | $10.00 |
| Audio | $32.00 | $64.00 |
Audio processing is substantially more expensive than text processing under these rates. Applications that send large amounts of audio or generate lengthy spoken responses should therefore monitor both audio input and audio output usage separately. A short text instruction paired with a long audio exchange can have a very different cost profile from a text-only request.
These are token-based API prices, not a flat subscription price. The final cost of an interaction depends on the amount and type of input and output consumed. Because the model is scheduled for shutdown, teams should also include migration work in the total cost of adopting it for a new project.
Main strengths and trade-offs
Direct audio conversation
GPT-Audio’s clearest strength is that audio is part of the model interaction rather than an afterthought. An application can design a spoken exchange around a model that accepts audio and produces audio, reducing the need to treat speech recognition, text reasoning, and speech synthesis as entirely separate user-facing stages.
Streaming and function calling
Streaming makes the model more suitable for interactive voice interfaces than a workflow that waits for a complete response before showing or playing anything. Function calling extends the model beyond conversation by allowing it to interact with application-controlled tools. Together, these features support assistants that can both speak and take authorized steps in a software system.
Important limitations
GPT-Audio does not support image or video input, so it is not appropriate for applications that need a single model to interpret visual media alongside speech. It also does not support structured outputs. If an application requires reliably machine-readable JSON for downstream processing, this model is a poor fit according to the supplied documentation, even though function calling is available.
Fine-tuning is not supported, which limits the ability to customize the model through provider-managed training. The model also has a relatively high audio price compared with its listed text price, making long audio sessions more expensive than equivalent text interactions.
Finally, lifecycle status is a decisive limitation. GPT-Audio is deprecated and scheduled for shutdown. A technically suitable model that cannot remain available for the required life of a product is not a suitable long-term foundation.
Reasoning, coding, speed, and cost profile
The supplied comparative assessment gives GPT-Audio a reasoning score of 5 out of 10, a coding score of 2 out of 10, a speed score of 7 out of 10, and a cost score of 5 out of 10. These are editorial estimates, not scores published by OpenAI, and they should be treated as directional rather than benchmark results.
The estimates suggest that GPT-Audio is best understood as an interaction-focused audio model rather than a specialist for difficult reasoning or software development. Its relatively stronger speed assessment is relevant to conversational voice applications, where response delay affects the user experience. Its low coding assessment means teams should not select it primarily for code generation or software-engineering work. Its middle cost assessment also needs to be interpreted alongside the much higher listed rates for audio tokens.
For a voice assistant, the practical trade-off is often responsiveness and direct speech handling versus cost, customization, and depth in non-audio tasks. GPT-Audio’s value comes from combining audio understanding and spoken output in one model interaction, not from being the strongest option for every reasoning, coding, or media task.
Best use cases for GPT-Audio
- Audio-enabled chat applications: Products can let users speak requests and receive spoken answers, with text available where useful.
- Voice interfaces: GPT-Audio can serve as the conversational model behind hands-free or speech-led product interactions.
- Spoken assistants: Streaming and function calling make it possible to combine spoken responses with application-controlled actions.
- Audio understanding and generation: Applications that need the model to interpret audio and answer with audio can use its core modality pairing directly.
- Existing Chat Completions integrations: Teams maintaining a current GPT-Audio deployment may use the documented model while preparing a migration.
These use cases are most defensible when the application genuinely needs direct audio input and output. If audio is only converted to text before processing and the response is later converted to speech by another system, a different architecture may offer better control or lifecycle stability.
When to choose GPT-Audio—and when not to
Choose GPT-Audio when an existing OpenAI Chat Completions application needs a model that can understand audio and generate spoken responses, and when the integration can be migrated before the announced shutdown. Its combination of audio modalities, streaming, and function calling is useful for prototypes, evaluations, and maintained systems that specifically depend on this model’s behavior.
Do not choose GPT-Audio for a new long-lived production deployment unless there is a specific, temporary reason to do so. OpenAI recommends GPT-Audio-1.5 as the replacement, so new projects should evaluate that successor rather than building around a deprecated model. The supplied research does not provide GPT-Audio-1.5’s specifications, pricing, or feature compatibility, so those details must be verified separately before migration.
Another option may be more appropriate in several situations:
- Use a visual-capable model when image or video understanding is required; GPT-Audio does not support those inputs.
- Use a structured-output-capable model when downstream software requires validated JSON rather than conversational text or audio.
- Use a text-focused model when the workload is primarily reasoning, writing, or coding and does not need direct audio interaction.
- Use a different voice architecture when the application requires advanced real-time voice-agent behavior, because GPT-Audio is not positioned in the supplied research as the preferred option for that category.
- Use a separate or specialized transcription workflow when transcription alone is the requirement rather than two-way audio conversation.
Bottom line
GPT-Audio is a focused OpenAI model for two-way audio interaction through Chat Completions. It accepts text and audio, returns text and spoken audio, supports streaming and function calling, and offers a 128,000-token context window. Those capabilities make it a reasonable fit for voice interfaces and audio-enabled assistants that need more than text-only conversation.
Its limitations are equally important: no image or video input, no fine-tuning, no structured outputs, relatively high audio-token pricing, and only modest editorial assessments for reasoning and coding. Most importantly, OpenAI has deprecated the model and scheduled its shutdown for January 20, 2027. That makes GPT-Audio primarily an existing-integration or transition target, while new long-term deployments should investigate the recommended GPT-Audio-1.5 replacement.

