What is GPT-Audio Mini?
GPT-Audio Mini is OpenAI's compact audio-capable model for applications that need conversations in text, speech, or a combination of both. It accepts text and audio input and can return text and audio output through the Chat Completions API. The model is positioned as a lower-cost alternative within OpenAI's GPT-Audio family, making it relevant to developers who need audio interaction but do not want to use the more expensive GPT Audio model.
The canonical model ID is gpt-audio-mini. OpenAI released the model on October 6, 2025. The alias currently points to the gpt-audio-mini-2025-12-15 snapshot, while gpt-audio-mini-2025-10-06 is an older snapshot.
Where GPT-Audio Mini fits in OpenAI's lineup
GPT-Audio Mini is designed around an audio conversation use case rather than general-purpose visual reasoning or high-end, continuously streamed voice interaction. Its defining advantage is the combination of audio input and audio output with a lower text-token price than the larger GPT Audio option. That positioning makes it suitable for turn-based voice assistants, audio-enabled chat, and early-stage customer-support experiences.
There is an important qualification for new projects: OpenAI lists GPT-Audio Mini as deprecated and has announced that the model family will be removed from the API on January 20, 2027. OpenAI identifies GPT-Audio 1.5 as the recommended replacement. The model may still be accessible before that date, but a new long-lived production system should account for the planned shutdown rather than treating GPT-Audio Mini as a stable long-term choice.
Capabilities and supported modalities
GPT-Audio Mini supports both text and audio at the input and output stages. In practical terms, an application can send written instructions or audio content, then receive a written response, an audio response, or both where the API workflow supports that format.
- Text input: Supported.
- Audio input: Supported.
- Text output: Supported.
- Audio output: Supported.
- Image input: Not supported.
- Video input: Not supported.
- Function calling: Supported.
- Batch API: Supported.
Function calling allows the model to request an application-defined function, such as looking up an account record or creating a support ticket. The application remains responsible for executing that function and returning the result. This makes GPT-Audio Mini more useful than a voice-only response generator for workflows that need to take action, provided the interaction does not depend on continuous streaming.
Context window and output limit
The model has a 128,000-token context window. The context window is the amount of model-visible conversation and other supplied text that can be included in a request, subject to the API's handling of audio and tokenization. A large window can help when a conversation includes substantial history, instructions, or retrieved application data, although longer requests may increase usage costs and may not be necessary for short voice exchanges.
GPT-Audio Mini supports a maximum output of 16,384 tokens. This is a ceiling rather than a requirement: most conversational replies should be much shorter. The limit is relevant for applications that generate long transcripts, detailed text responses, or extended tool-oriented output.
GPT-Audio Mini pricing
OpenAI lists text-token pricing of $0.60 per 1 million input tokens and $2.40 per 1 million output tokens. These are the supplied text-token rates. Audio usage may be represented through audio tokens, so a cost estimate based only on visible text can understate the cost of a voice application.
For budgeting, developers should consider the length of the conversation, the amount of audio exchanged, the number of turns, and whether conversation history is repeatedly sent with later requests. GPT-Audio Mini's lower text-token rates make it attractive for cost-sensitive prototypes and moderate-volume audio workflows, but the actual bill depends on the complete token usage pattern rather than the model name alone.
Main strengths and limitations
The clearest strength of GPT-Audio Mini is its balance of audio support and cost. It can handle audio in both directions while also supporting ordinary text interactions and function calling. Its 128,000-token context window is ample for many conversational and tool-assisted applications, and its 16,384-token output limit leaves room for responses longer than a typical spoken turn.
The principal limitations are functional and operational:
- No streaming: The supplied documentation does not list streaming support, so the model is better suited to turn-based exchanges than to low-latency, continuously flowing voice sessions.
- No visual input: It cannot accept images or video, making it unsuitable for workflows that require interpreting screens, photographs, documents as images, or video footage.
- No structured outputs: It is not listed as supporting structured outputs, so applications requiring provider-enforced JSON schemas should choose a model with that capability.
- No fine-tuning: The model cannot be fine-tuned according to the supplied specifications.
- Scheduled removal: Its announced January 20, 2027 shutdown makes migration planning necessary.
Reasoning, coding, speed, and cost profile
The supplied research assigns GPT-Audio Mini an editorial reasoning score of 5 out of 10 and a coding score of 4 out of 10. These are comparative editorial estimates, not benchmarks published by OpenAI. They suggest that the model should be evaluated primarily for audio conversation and task orchestration rather than selected as a specialist for difficult reasoning or complex software development.
The same editorial assessment gives the model a speed score of 8 out of 10 and a cost score of 8 out of 10. These scores are also subjective comparisons, not provider guarantees. They describe the intended trade-off: GPT-Audio Mini emphasizes relatively fast, economical interaction, while giving up capabilities such as streaming, structured outputs, visual input, and the broader use cases of more capable models.
Best use cases
GPT-Audio Mini is a reasonable fit when the application needs audio input and output, does not require continuous streaming, and benefits from lower-cost processing. Suitable examples include:
- Turn-based voice assistants that wait for a user to finish speaking before responding.
- Audio-enabled chat applications with text as a fallback or parallel output.
- Spoken customer-support prototypes that use function calls to retrieve information or initiate actions.
- Internal voice tools where the conversation history is substantial but visual understanding is unnecessary.
- Experiments with audio workflows where a scheduled migration to a newer model is acceptable.
For these scenarios, the model's function-calling support can connect an audio conversation to application services, while its text output can provide logs, captions, or a readable record of the exchange.
When another option may be more appropriate
Choose another model or plan a different architecture if the application requires real-time streaming voice interaction. GPT-Audio Mini is not listed as supporting streaming, so it is not a strong fit for a continuously responsive voice interface where partial audio or incremental responses are central to the user experience.
A model with image or video input is more appropriate for visual assistants, image analysis, or workflows that combine speech with screen or camera content. A model with structured-output support is preferable when downstream software requires responses to conform to a defined JSON schema. Fine-tuning is also unavailable, so applications that depend on training a model for a specialized behavior should look elsewhere.
Finally, GPT-Audio Mini is a poor default for a new system expected to operate beyond January 20, 2027. Its low cost and audio capabilities may still be useful for short-lived projects or migration-stage experiments, but OpenAI's announced replacement, GPT-Audio 1.5, should be considered for new production deployments when the project requires a longer expected service life.
Bottom line
GPT-Audio Mini is a lower-cost OpenAI audio model with text and audio input and output, function calling, a 128,000-token context window, and a 16,384-token maximum output. Its strongest use case is a cost-sensitive, turn-based audio application that does not need streaming, visual input, structured outputs, or fine-tuning. Because OpenAI has deprecated the model and scheduled its API removal for January 20, 2027, its capabilities should be weighed against migration work before it is used in a new long-term deployment.

