GPT-Audio

GPT-Audio

by OpenAI · Deprecated; scheduled for shutdown on January 20, 2027

GPT-Audio is OpenAI’s Chat Completions model for direct audio understanding and spoken responses. It supports text and audio input and output, streaming, function calling, a 128,000-token context window, and up to 16,384 output tokens. Text and audio are priced separately, with audio costing more. The model does not support images, video, fine-tuning, or structured outputs, and OpenAI has scheduled its shutdown for January 20, 2027.

Text Speech Reasoning Coding
GPT-Audio is OpenAI’s audio-capable model for Chat Completions. It is designed for voice interfaces, spoken assistants, and applications that need a model to interpret audio and respond with generated speech rather than relying only on text. The model supports both text and audio input and output, along with streaming and function calling, but it does not process images or video. Because OpenAI has deprecated GPT-Audio and scheduled its shutdown for January 20, 2027, it is mainly relevant for understanding existing integrations or evaluating short-lived deployments rather than starting a new long-term production system.
Outputs

What GPT-Audio can produce

Text Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Tool use Streaming Multimodal output
Model profile

Performance characteristics

5/10 Reasoning
2/10 Coding
7/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family GPT-Audio
Model type Multimodal
Context window 128K tokens
Maximum output 16K tokens
Knowledge cutoff 2023-10-01
Release date 2025-08-28
Status Deprecated; scheduled for shutdown on January 20, 2027
Deprecation date 2026-07-20
Shutdown date 2027-01-20
Knowledge cutoff notes

The official model page lists October 1, 2023 as the model's knowledge cutoff. Audio input, external context, or application-level web retrieval does not change this underlying cutoff.

Model notes

The canonical model ID is gpt-audio. Its documented snapshot is gpt-audio-2025-08-28. The model is OpenAI's first generally available audio model and supports audio and text input/output through Chat Completions. Function calling and streaming are supported. Structured outputs are not supported. OpenAI recommends migrating to gpt-audio-1.5 before the January 20, 2027 shutdown date. Reasoning, coding, speed, and cost values are editorial comparative estimates rather than vendor-published scores.

Cost

Model pricing

Input $2.50 per 1M text tokens; $32.00 per 1M audio tokens
Output $10.00 per 1M text tokens; $64.00 per 1M audio tokens
Model guide

GPT-Audio: OpenAI’s Audio-First Chat Completions Model

GPT-Audio is OpenAI’s generally available model for applications that need to understand audio and produce spoken responses through the Chat Completions API. It accepts text and audio inputs, returns text and audio outputs, supports streaming and function calling, and provides a 128,000-token context window. Its main practical limitation is lifecycle status: the model is deprecated, has a documented deprecation date of July 20, 2026, and is scheduled to shut down on January 20, 2027. OpenAI recommends GPT-Audio-1.5 as the replacement.

What is GPT-Audio?

GPT-Audio is an OpenAI multimodal model built for audio conversations in the Chat Completions API. In practical terms, an application can send the model written text, recorded or streamed audio, or a combination of text and audio, and receive either text or generated audio in response. This makes it suitable for conversational experiences where users speak naturally and the application answers aloud.

The model is not a general-purpose image-and-video model. Its documented modalities are text and audio: it can accept text and audio inputs and produce text and audio outputs. OpenAI describes it as its first generally available audio model, with support for direct audio understanding and generation through Chat Completions.

The canonical model identifier is gpt-audio. OpenAI also documents the snapshot identifier gpt-audio-2025-08-28. The model was released on August 28, 2025, according to the supplied model data.

Where GPT-Audio fits in OpenAI’s lineup

GPT-Audio occupies a specialized position rather than serving as a universal model for every media type. Its purpose is direct audio interaction inside the Chat Completions workflow. That distinguishes it from text-only models, which require a separate speech-recognition step before the language model can process a spoken request, and from models intended for advanced real-time voice-agent workflows.

That position is now transitional. OpenAI lists GPT-Audio as deprecated, with a documented deprecation date of July 20, 2026, and a shutdown date of January 20, 2027. OpenAI recommends migrating to GPT-Audio-1.5 before the shutdown. GPT-Audio can therefore still be relevant when maintaining an existing integration or testing behavior tied specifically to this model, but it is not the safest choice for a new application expected to operate beyond the shutdown date.

Core capabilities and supported modalities

GPT-Audio’s defining capability is native handling of both audio and text within a conversational model request. Audio input can represent a spoken user request or other audio content the application needs the model to interpret. Audio output allows the response to be delivered as speech rather than only as text.

CapabilityGPT-Audio support
Text inputSupported
Audio inputSupported
Text outputSupported
Audio or speech outputSupported
Image inputNot supported
Video inputNot supported
StreamingSupported
Function callingSupported
Fine-tuningNot supported
Structured outputNot supported

Streaming is important for voice experiences because it can allow an application to begin handling or delivering a response before the entire exchange is complete. Function calling allows the model to request application-defined operations, such as looking up information or triggering a workflow. The model itself does not automatically perform arbitrary actions; the surrounding application must implement and authorize the available functions.

Context window and output limit

GPT-Audio has a documented context length of 128,000 tokens. A token is a unit of text or model-processing data, so the context window represents the amount of conversation and other supported content the model can consider within a request. A larger context can help when an audio conversation includes substantial accompanying text or a long prior exchange, although the usable capacity depends on the complete request and the way audio is represented.

The maximum output limit is 16,384 tokens. This is a ceiling on the generated response rather than a promise that every request will produce an output of that size. In a spoken interface, practical response length should usually be controlled by the application so that answers remain timely and easy to follow.

The model’s listed knowledge cutoff is October 1, 2023. Sending audio, adding external context, or connecting application-level retrieval does not change that underlying cutoff. If an application needs current information, it must supply that information through its own tools or external systems; GPT-Audio is not documented as supporting web search directly.

GPT-Audio pricing

GPT-Audio uses separate token rates for text and audio, with different prices for input and output. The supplied OpenAI pricing information is:

Token typeInput price per 1 million tokensOutput price per 1 million tokens
Text$2.50$10.00
Audio$32.00$64.00

Audio processing is substantially more expensive than text processing under these rates. Applications that send large amounts of audio or generate lengthy spoken responses should therefore monitor both audio input and audio output usage separately. A short text instruction paired with a long audio exchange can have a very different cost profile from a text-only request.

These are token-based API prices, not a flat subscription price. The final cost of an interaction depends on the amount and type of input and output consumed. Because the model is scheduled for shutdown, teams should also include migration work in the total cost of adopting it for a new project.

Main strengths and trade-offs

Direct audio conversation

GPT-Audio’s clearest strength is that audio is part of the model interaction rather than an afterthought. An application can design a spoken exchange around a model that accepts audio and produces audio, reducing the need to treat speech recognition, text reasoning, and speech synthesis as entirely separate user-facing stages.

Streaming and function calling

Streaming makes the model more suitable for interactive voice interfaces than a workflow that waits for a complete response before showing or playing anything. Function calling extends the model beyond conversation by allowing it to interact with application-controlled tools. Together, these features support assistants that can both speak and take authorized steps in a software system.

Important limitations

GPT-Audio does not support image or video input, so it is not appropriate for applications that need a single model to interpret visual media alongside speech. It also does not support structured outputs. If an application requires reliably machine-readable JSON for downstream processing, this model is a poor fit according to the supplied documentation, even though function calling is available.

Fine-tuning is not supported, which limits the ability to customize the model through provider-managed training. The model also has a relatively high audio price compared with its listed text price, making long audio sessions more expensive than equivalent text interactions.

Finally, lifecycle status is a decisive limitation. GPT-Audio is deprecated and scheduled for shutdown. A technically suitable model that cannot remain available for the required life of a product is not a suitable long-term foundation.

Reasoning, coding, speed, and cost profile

The supplied comparative assessment gives GPT-Audio a reasoning score of 5 out of 10, a coding score of 2 out of 10, a speed score of 7 out of 10, and a cost score of 5 out of 10. These are editorial estimates, not scores published by OpenAI, and they should be treated as directional rather than benchmark results.

The estimates suggest that GPT-Audio is best understood as an interaction-focused audio model rather than a specialist for difficult reasoning or software development. Its relatively stronger speed assessment is relevant to conversational voice applications, where response delay affects the user experience. Its low coding assessment means teams should not select it primarily for code generation or software-engineering work. Its middle cost assessment also needs to be interpreted alongside the much higher listed rates for audio tokens.

For a voice assistant, the practical trade-off is often responsiveness and direct speech handling versus cost, customization, and depth in non-audio tasks. GPT-Audio’s value comes from combining audio understanding and spoken output in one model interaction, not from being the strongest option for every reasoning, coding, or media task.

Best use cases for GPT-Audio

  • Audio-enabled chat applications: Products can let users speak requests and receive spoken answers, with text available where useful.
  • Voice interfaces: GPT-Audio can serve as the conversational model behind hands-free or speech-led product interactions.
  • Spoken assistants: Streaming and function calling make it possible to combine spoken responses with application-controlled actions.
  • Audio understanding and generation: Applications that need the model to interpret audio and answer with audio can use its core modality pairing directly.
  • Existing Chat Completions integrations: Teams maintaining a current GPT-Audio deployment may use the documented model while preparing a migration.

These use cases are most defensible when the application genuinely needs direct audio input and output. If audio is only converted to text before processing and the response is later converted to speech by another system, a different architecture may offer better control or lifecycle stability.

When to choose GPT-Audio—and when not to

Choose GPT-Audio when an existing OpenAI Chat Completions application needs a model that can understand audio and generate spoken responses, and when the integration can be migrated before the announced shutdown. Its combination of audio modalities, streaming, and function calling is useful for prototypes, evaluations, and maintained systems that specifically depend on this model’s behavior.

Do not choose GPT-Audio for a new long-lived production deployment unless there is a specific, temporary reason to do so. OpenAI recommends GPT-Audio-1.5 as the replacement, so new projects should evaluate that successor rather than building around a deprecated model. The supplied research does not provide GPT-Audio-1.5’s specifications, pricing, or feature compatibility, so those details must be verified separately before migration.

Another option may be more appropriate in several situations:

  • Use a visual-capable model when image or video understanding is required; GPT-Audio does not support those inputs.
  • Use a structured-output-capable model when downstream software requires validated JSON rather than conversational text or audio.
  • Use a text-focused model when the workload is primarily reasoning, writing, or coding and does not need direct audio interaction.
  • Use a different voice architecture when the application requires advanced real-time voice-agent behavior, because GPT-Audio is not positioned in the supplied research as the preferred option for that category.
  • Use a separate or specialized transcription workflow when transcription alone is the requirement rather than two-way audio conversation.

Bottom line

GPT-Audio is a focused OpenAI model for two-way audio interaction through Chat Completions. It accepts text and audio, returns text and spoken audio, supports streaming and function calling, and offers a 128,000-token context window. Those capabilities make it a reasonable fit for voice interfaces and audio-enabled assistants that need more than text-only conversation.

Its limitations are equally important: no image or video input, no fine-tuning, no structured outputs, relatively high audio-token pricing, and only modest editorial assessments for reasoning and coding. Most importantly, OpenAI has deprecated the model and scheduled its shutdown for January 20, 2027. That makes GPT-Audio primarily an existing-integration or transition target, while new long-term deployments should investigate the recommended GPT-Audio-1.5 replacement.


Answers to Frequently Asked Questions

When should developers choose GPT-Audio?
GPT-Audio is mainly appropriate for existing Chat Completions integrations, prototypes, and temporary voice applications that need direct audio input and spoken output. It is not recommended as the foundation for a new long-term production system because it is deprecated and scheduled for shutdown.
Is GPT-Audio still available, and when will it shut down?
GPT-Audio is deprecated. OpenAI has documented a deprecation date of July 20, 2026, and a shutdown date of January 20, 2027. OpenAI recommends migrating to GPT-Audio-1.5 before the shutdown.
How much does GPT-Audio cost?
GPT-Audio is priced by token type and direction. Text input costs $2.50 per 1 million tokens and text output costs $10.00. Audio input costs $32.00 per 1 million tokens and audio output costs $64.00. Actual costs depend on the amount of text and audio processed.
What are the supported modalities and features of GPT-Audio?
GPT-Audio supports text input, audio input, text output, and audio output. It also supports streaming and function calling. It does not support image or video input, fine-tuning, or structured outputs.
What is GPT-Audio and what can it do?
GPT-Audio is an OpenAI multimodal model for audio conversations in the Chat Completions API. It accepts text and audio inputs and can return text or generated audio, making it suitable for voice interfaces, spoken assistants, and audio-enabled chat applications.


Sources 5
Provider

About OpenAI