What GPT-4o Mini Audio is
GPT-4o Mini Audio was OpenAI's smaller preview model for applications that needed to understand and generate audio. Its canonical API identifier was gpt-4o-mini-audio-preview, and OpenAI also documented a dated snapshot called gpt-4o-mini-audio-preview-2024-12-17.
The model was intended for developers building voice assistants, conversational audio interfaces, audio-understanding tools, and applications that needed spoken responses rather than text alone. It accepted both text and audio input and could produce either text or audio output. That combination made it different from a text-only language model, while its smaller positioning was aimed at reducing cost for suitable audio workloads.
GPT-4o Mini Audio is not a current, recommended model for new deployments. OpenAI lists it as deprecated and has announced that the gpt-4o-mini-audio family will be removed from the API on January 20, 2027. Developers maintaining an existing integration should plan a migration rather than treating this model as a long-term foundation.
Audio and text capabilities
The model supported the following input and output combinations:
| Capability | Support |
|---|---|
| Text input | Yes |
| Audio input | Yes |
| Text output | Yes |
| Audio output | Yes |
| Image input or output | No |
| Video input or output | No |
In practical terms, an application could send a spoken request and receive either a written answer or spoken audio. It could also use text as the input while requesting an audio response. The model therefore suited conversational voice experiences, but it was not a general image-and-video multimodal model.
OpenAI documented streaming support, which allowed partial results to be delivered while a response was being generated instead of waiting for the complete response. This was particularly relevant to voice interfaces, where waiting for an entire answer before playback can make a conversation feel slow. The model also supported function calling, allowing an application to request actions or retrieve information through developer-defined functions. Function calling did not make the model an autonomous system; the surrounding application still had to execute and validate those functions.
API endpoints and integration scope
OpenAI listed GPT-4o Mini Audio for a broad set of API surfaces, including Chat Completions, Responses, Realtime, Live sessions, Realtime translation, Realtime transcription sessions, Assistants, and Batch. The exact behavior of audio generation and streaming could vary by endpoint, so an existing integration should be checked against the endpoint-specific documentation before migration or maintenance work.
The model's documented feature set did not include Structured Outputs or fine-tuning. Structured Outputs are designed to constrain a response to a specified schema; their absence matters for applications that require reliably formatted machine-readable results. Developers could still use ordinary text responses and application-side validation where appropriate, but that should not be confused with native Structured Outputs support.
Context window and output limit
GPT-4o Mini Audio had a 128,000-token context window and a maximum output of 16,384 tokens. A context window is the amount of material the model can consider in one request, including the conversation and other supplied content. The limit was large enough for extended conversations or substantial text context, but it did not remove the need to manage long-running audio sessions and conversation history carefully.
The maximum output figure is a token limit, not a guaranteed duration of spoken audio. Actual audio response length depends on the request, selected output, and endpoint behavior. The model's listed knowledge cutoff was October 1, 2023, so it should not be treated as having built-in knowledge of events after that date unless an application supplied current information through another mechanism.
GPT-4o Mini Audio pricing
OpenAI documented separate rates for text tokens and audio tokens. The supplied preview pricing was:
| Usage type | Price per 1 million tokens |
|---|---|
| Text input | $0.15 |
| Text output | $0.60 |
| Audio input | $10.00 |
| Audio output | $20.00 |
These are API token prices, not consumer ChatGPT subscription prices. The important cost distinction is between text and audio: audio input and output were substantially more expensive per million tokens than text input and output. An application that converts every interaction into audio may therefore have a very different cost profile from one that uses audio only for the user's speech and returns text.
Because the model is deprecated, developers should verify the currently applicable pricing and endpoint availability before making cost projections for a migration. The figures above describe the documented GPT-4o Mini Audio preview pricing and should not be assumed to apply to GPT-Audio-1.5 or another replacement.
Strengths and trade-offs
The model's main strength was its focus on relatively economical audio-capable interactions. Compared with using a larger audio model for every request, its documented positioning offered a lower-cost option for straightforward voice interfaces and audio-aware applications. It also combined audio input, audio output, streaming, and function calling in one model family, which could simplify the design of conversational applications.
Its limitations were equally important. It was a preview model, lacked image and video support, did not support Structured Outputs or fine-tuning, and had a knowledge cutoff in 2023. The model was also deprecated, so lifecycle risk now outweighs its original cost advantage for most new projects. A low per-token price is not enough to justify new integration work when the provider has already announced a removal date.
The research record does not provide provider-published benchmark scores for reasoning, coding, speed, or overall quality. Editorial comparative estimates rated its reasoning capability at 3 out of 10, coding at 4 out of 10, speed at 8 out of 10, and cost at 8 out of 10. These are subjective editorial assessments, not OpenAI benchmarks. They suggest a model aimed at fast, economical audio interactions rather than advanced reasoning or demanding software-development tasks, but they should not be interpreted as formal performance measurements.
When to choose GPT-4o Mini Audio
For a new application, GPT-4o Mini Audio is generally difficult to justify because of its deprecated status and announced shutdown. It may still be relevant when investigating or maintaining an existing integration that already depends on its specific API behavior. In that situation, its audio input and output, streaming, and function-calling support may explain why it was originally selected.
Its historical use cases included:
- Voice assistants that needed spoken input and spoken responses.
- Conversational audio interfaces with streaming interaction.
- Applications that converted audio conversations into text or generated audio from text.
- Lower-cost prototypes where advanced reasoning, image processing, or structured response generation was not required.
- Audio-aware workflows that used function calling to connect a conversation to application actions.
Another option is more appropriate when the project requires a supported model with a longer lifecycle, native structured response support, fine-tuning, image or video processing, advanced reasoning, or a current pricing policy. OpenAI recommends GPT-Audio-1.5 as the migration target. That recommendation should be evaluated against the application's endpoint, latency, audio quality, function-calling behavior, and total cost rather than assumed to be a drop-in replacement.
Availability and migration considerations
OpenAI announced that the gpt-4o-mini-audio family will be removed from the API on January 20, 2027. Developers should inventory model identifiers, endpoints, audio formats, streaming behavior, function definitions, error handling, and token accounting before moving an application. The dated snapshot identifier should not be assumed to remain available simply because an older integration still references it.
A sensible migration process is to test the replacement with representative conversations, including short requests, long context, interruptions, function calls, and audio responses. Compare not only response quality but also latency and the separate costs of audio input, audio output, text input, and text output. If an application does not actually need spoken output, a text-oriented design may also have a different cost and integration profile.
Bottom line
GPT-4o Mini Audio was a compact OpenAI preview model for text-and-audio conversations. Its combination of audio input and output, streaming, function calling, a 128,000-token context window, and lower text-token pricing made it useful for an earlier generation of voice applications. Today, however, its deprecated status and January 20, 2027 removal date are the defining practical facts. It is mainly relevant for understanding or migrating existing deployments, while new projects should evaluate the provider's recommended successor or another currently supported audio-capable option.

