GPT-4o

GPT-4o Mini Audio

by OpenAI · Deprecated; scheduled for API shutdown on 2027-01-20

GPT-4o Mini Audio was a smaller OpenAI preview model for text-and-audio applications, offering audio input and output, streaming, function calling, a 128,000-token context window, and separate text and audio token pricing. It is deprecated and scheduled for API removal on January 20, 2027.

Text Speech Reasoning Coding
GPT-4o Mini Audio was designed for lower-cost audio understanding and generation compared with GPT-4o Audio. It could accept text or audio and return text or audio, making it suitable for conversational voice interfaces and other audio-aware applications. However, it is no longer a good choice for new production deployments: OpenAI lists the model as deprecated and recommends GPT-Audio-1.5 as the migration target.
Outputs

What GPT-4o Mini Audio can produce

Text Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Tool use Streaming Batch API Multimodal output
Model profile

Performance characteristics

3/10 Reasoning
4/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o
Model type Multimodal
Context window 128K tokens
Maximum output 16K tokens
Knowledge cutoff 2023-10-01
Release date 2024-12-17
Status Deprecated; scheduled for API shutdown on 2027-01-20
Deprecation date 2026-07-20
Shutdown date 2027-01-20
Knowledge cutoff notes

The official model page lists October 1, 2023 as the model's knowledge cutoff.

Model notes

The canonical model identifier is gpt-4o-mini-audio-preview. The dated snapshot gpt-4o-mini-audio-preview-2024-12-17 is listed as deprecated. OpenAI announced that the gpt-4o-mini-audio family will be removed from the API on January 20, 2027, and recommends GPT-Audio-1.5 as the replacement. The model page lists Chat Completions, Responses, Realtime, Live sessions, Realtime translation, Realtime transcription sessions, Assistants, and Batch endpoints. Text and audio pricing use different token rates. Structured Outputs and fine-tuning are documented as unsupported. Editorial scores are comparative estimates rather than provider benchmarks.

Cost

Model pricing

Input Text: $0.15 per 1M tokens; audio: $10.00 per 1M tokens
Output Text: $0.60 per 1M tokens; audio: $20.00 per 1M tokens
Model guide

GPT-4o Mini Audio: A Deprecated, Lower-Cost Model for Voice Applications

GPT-4o Mini Audio was a smaller OpenAI preview model for applications that needed both audio input and audio output through the API. It accepted text and audio, returned text or spoken audio, supported streaming and function calling, and offered lower text-token prices than larger audio-capable models. OpenAI has deprecated it and scheduled the gpt-4o-mini-audio family for API removal on January 20, 2027.

What GPT-4o Mini Audio is

GPT-4o Mini Audio was OpenAI's smaller preview model for applications that needed to understand and generate audio. Its canonical API identifier was gpt-4o-mini-audio-preview, and OpenAI also documented a dated snapshot called gpt-4o-mini-audio-preview-2024-12-17.

The model was intended for developers building voice assistants, conversational audio interfaces, audio-understanding tools, and applications that needed spoken responses rather than text alone. It accepted both text and audio input and could produce either text or audio output. That combination made it different from a text-only language model, while its smaller positioning was aimed at reducing cost for suitable audio workloads.

GPT-4o Mini Audio is not a current, recommended model for new deployments. OpenAI lists it as deprecated and has announced that the gpt-4o-mini-audio family will be removed from the API on January 20, 2027. Developers maintaining an existing integration should plan a migration rather than treating this model as a long-term foundation.

Audio and text capabilities

The model supported the following input and output combinations:

CapabilitySupport
Text inputYes
Audio inputYes
Text outputYes
Audio outputYes
Image input or outputNo
Video input or outputNo

In practical terms, an application could send a spoken request and receive either a written answer or spoken audio. It could also use text as the input while requesting an audio response. The model therefore suited conversational voice experiences, but it was not a general image-and-video multimodal model.

OpenAI documented streaming support, which allowed partial results to be delivered while a response was being generated instead of waiting for the complete response. This was particularly relevant to voice interfaces, where waiting for an entire answer before playback can make a conversation feel slow. The model also supported function calling, allowing an application to request actions or retrieve information through developer-defined functions. Function calling did not make the model an autonomous system; the surrounding application still had to execute and validate those functions.

API endpoints and integration scope

OpenAI listed GPT-4o Mini Audio for a broad set of API surfaces, including Chat Completions, Responses, Realtime, Live sessions, Realtime translation, Realtime transcription sessions, Assistants, and Batch. The exact behavior of audio generation and streaming could vary by endpoint, so an existing integration should be checked against the endpoint-specific documentation before migration or maintenance work.

The model's documented feature set did not include Structured Outputs or fine-tuning. Structured Outputs are designed to constrain a response to a specified schema; their absence matters for applications that require reliably formatted machine-readable results. Developers could still use ordinary text responses and application-side validation where appropriate, but that should not be confused with native Structured Outputs support.

Context window and output limit

GPT-4o Mini Audio had a 128,000-token context window and a maximum output of 16,384 tokens. A context window is the amount of material the model can consider in one request, including the conversation and other supplied content. The limit was large enough for extended conversations or substantial text context, but it did not remove the need to manage long-running audio sessions and conversation history carefully.

The maximum output figure is a token limit, not a guaranteed duration of spoken audio. Actual audio response length depends on the request, selected output, and endpoint behavior. The model's listed knowledge cutoff was October 1, 2023, so it should not be treated as having built-in knowledge of events after that date unless an application supplied current information through another mechanism.

GPT-4o Mini Audio pricing

OpenAI documented separate rates for text tokens and audio tokens. The supplied preview pricing was:

Usage typePrice per 1 million tokens
Text input$0.15
Text output$0.60
Audio input$10.00
Audio output$20.00

These are API token prices, not consumer ChatGPT subscription prices. The important cost distinction is between text and audio: audio input and output were substantially more expensive per million tokens than text input and output. An application that converts every interaction into audio may therefore have a very different cost profile from one that uses audio only for the user's speech and returns text.

Because the model is deprecated, developers should verify the currently applicable pricing and endpoint availability before making cost projections for a migration. The figures above describe the documented GPT-4o Mini Audio preview pricing and should not be assumed to apply to GPT-Audio-1.5 or another replacement.

Strengths and trade-offs

The model's main strength was its focus on relatively economical audio-capable interactions. Compared with using a larger audio model for every request, its documented positioning offered a lower-cost option for straightforward voice interfaces and audio-aware applications. It also combined audio input, audio output, streaming, and function calling in one model family, which could simplify the design of conversational applications.

Its limitations were equally important. It was a preview model, lacked image and video support, did not support Structured Outputs or fine-tuning, and had a knowledge cutoff in 2023. The model was also deprecated, so lifecycle risk now outweighs its original cost advantage for most new projects. A low per-token price is not enough to justify new integration work when the provider has already announced a removal date.

The research record does not provide provider-published benchmark scores for reasoning, coding, speed, or overall quality. Editorial comparative estimates rated its reasoning capability at 3 out of 10, coding at 4 out of 10, speed at 8 out of 10, and cost at 8 out of 10. These are subjective editorial assessments, not OpenAI benchmarks. They suggest a model aimed at fast, economical audio interactions rather than advanced reasoning or demanding software-development tasks, but they should not be interpreted as formal performance measurements.

When to choose GPT-4o Mini Audio

For a new application, GPT-4o Mini Audio is generally difficult to justify because of its deprecated status and announced shutdown. It may still be relevant when investigating or maintaining an existing integration that already depends on its specific API behavior. In that situation, its audio input and output, streaming, and function-calling support may explain why it was originally selected.

Its historical use cases included:

  • Voice assistants that needed spoken input and spoken responses.
  • Conversational audio interfaces with streaming interaction.
  • Applications that converted audio conversations into text or generated audio from text.
  • Lower-cost prototypes where advanced reasoning, image processing, or structured response generation was not required.
  • Audio-aware workflows that used function calling to connect a conversation to application actions.

Another option is more appropriate when the project requires a supported model with a longer lifecycle, native structured response support, fine-tuning, image or video processing, advanced reasoning, or a current pricing policy. OpenAI recommends GPT-Audio-1.5 as the migration target. That recommendation should be evaluated against the application's endpoint, latency, audio quality, function-calling behavior, and total cost rather than assumed to be a drop-in replacement.

Availability and migration considerations

OpenAI announced that the gpt-4o-mini-audio family will be removed from the API on January 20, 2027. Developers should inventory model identifiers, endpoints, audio formats, streaming behavior, function definitions, error handling, and token accounting before moving an application. The dated snapshot identifier should not be assumed to remain available simply because an older integration still references it.

A sensible migration process is to test the replacement with representative conversations, including short requests, long context, interruptions, function calls, and audio responses. Compare not only response quality but also latency and the separate costs of audio input, audio output, text input, and text output. If an application does not actually need spoken output, a text-oriented design may also have a different cost and integration profile.

Bottom line

GPT-4o Mini Audio was a compact OpenAI preview model for text-and-audio conversations. Its combination of audio input and output, streaming, function calling, a 128,000-token context window, and lower text-token pricing made it useful for an earlier generation of voice applications. Today, however, its deprecated status and January 20, 2027 removal date are the defining practical facts. It is mainly relevant for understanding or migrating existing deployments, while new projects should evaluate the provider's recommended successor or another currently supported audio-capable option.


Answers to Frequently Asked Questions

What should developers use instead of GPT-4o Mini Audio?
OpenAI recommends evaluating GPT-Audio-1.5 as a migration target. Developers should test the replacement with representative conversations and compare endpoint compatibility, audio quality, latency, function calling, token accounting, and total cost before switching.
How much did GPT-4o Mini Audio cost?
The documented preview pricing was $0.15 per 1 million text input tokens, $0.60 per 1 million text output tokens, $10.00 per 1 million audio input tokens, and $20.00 per 1 million audio output tokens. These historical rates should not be assumed to apply to replacement models.
Is GPT-4o Mini Audio still available for new applications?
No. GPT-4o Mini Audio is deprecated and is not recommended for new deployments. OpenAI announced that the gpt-4o-mini-audio family will be removed from the API on January 20, 2027.
What audio and text capabilities did GPT-4o Mini Audio support?
The model supported text and audio input as well as text and audio output. It also supported streaming and function calling, but it did not support image or video input and output, Structured Outputs, or fine-tuning.
What is GPT-4o Mini Audio?
GPT-4o Mini Audio was OpenAI’s smaller preview model for applications that needed to understand and generate audio. It accepted text and audio input and could return text or audio output, making it suitable for voice assistants and conversational audio interfaces.


Sources 4
Provider

About OpenAI