What is GPT-Realtime Mini?
GPT-Realtime Mini is an OpenAI realtime model for interactive applications, particularly those that communicate through voice. Unlike a conventional text-only model interaction, a realtime session can receive audio and return spoken audio with low conversational delay. It can also work with text and image inputs, allowing an application to combine spoken conversation with written instructions or visual context.
The model is positioned as a smaller, more cost-efficient member of OpenAI’s realtime model family. Its main appeal is not maximum reasoning performance; it is the balance between responsiveness, multimodal input, audio output, and operating cost. Typical applications include voice agents, customer-service interfaces, speech-to-speech assistants, and realtime interfaces that use function calling to access application actions or data.
There is an important lifecycle qualification: GPT-Realtime Mini is deprecated. OpenAI has announced API shutdown for January 20, 2027, and recommends GPT-Realtime-2.1 Mini as the replacement. It may still be useful for existing integrations or short-term evaluation, but the planned shutdown should be treated as a central factor in any production decision.
Supported input and output modalities
GPT-Realtime Mini accepts three input types: text, images, and audio. It produces text and audio output. This makes it suitable for both speech-to-speech conversations and applications where users speak but the interface displays a written answer as well.
- Text input: Supported.
- Image input: Supported for visual context, but the model does not generate images.
- Audio input: Supported for spoken interaction and realtime voice sessions.
- Video input: Not supported.
- Text output: Supported.
- Audio output: Supported, including spoken responses.
- Image and video output: Not supported.
For example, a voice assistant could receive a spoken question, inspect an image supplied by the user, call a business function, and return both a spoken response and text for the application interface. The supplied research does not indicate that GPT-Realtime Mini can directly process video streams, so video-based use cases require another design or model.
Context window and technical specifications
The model has a 32,000-token context window. A context window is the amount of conversation, instructions, tool information, and other model-visible content that can be considered within a session. This is large enough for many ordinary voice-agent interactions, although long-running sessions and applications that attach substantial documents or tool output will need to manage context carefully.
| Specification | GPT-Realtime Mini |
|---|---|
| Provider | OpenAI |
| Model family | GPT-Realtime |
| Model type | Realtime |
| Context window | 32,000 tokens |
| Maximum output | 4,096 tokens |
| Input | Text, image, and audio |
| Output | Text and audio |
| Function calling | Supported |
| Fine-tuning | Not supported |
| Structured outputs | Not supported |
The maximum output limit is 4,096 tokens. In a voice application, this limit applies to the model response rather than directly describing the duration of an audio clip; actual audio usage and response length depend on the session and modality. The model supports function calling, which allows an application to expose defined operations such as checking an order, booking an appointment, or retrieving account information. The model can request a function, but the surrounding application remains responsible for executing it and handling the result.
Realtime connections and API support
GPT-Realtime Mini is documented for OpenAI’s Realtime API and related session workflows. Connections can use WebRTC, WebSocket, or SIP. WebRTC is commonly associated with browser or client voice experiences, WebSocket with server-managed connections, and SIP with telephony integrations; the appropriate choice depends on the architecture of the application.
The model documentation lists support across several OpenAI endpoint categories, including Chat Completions, Responses, Live sessions, realtime translation, realtime transcription sessions, and Batch. The model should still be evaluated in the specific endpoint and session pattern an application needs, because a model being listed for an endpoint does not remove the practical differences between ordinary request-response use and a realtime audio session.
GPT-Realtime Mini supports function calling but not structured outputs. Structured outputs are useful when an application requires the model to conform to a formally defined response schema. Since that capability is not supported here, developers who need reliable schema-constrained responses may need additional validation, a different workflow, or another model.
Pricing and cost trade-offs
OpenAI’s listed text-token prices are:
- Text input: $0.60 per 1 million tokens.
- Cached text input: $0.06 per 1 million tokens.
- Text output: $2.40 per 1 million tokens.
These figures make GPT-Realtime Mini a cost-conscious option for text-token usage. Cached input is substantially cheaper than uncached input, which may matter when applications repeatedly send stable instructions or other reusable context. The supplied pricing information also warns that realtime sessions can incur modality-specific audio-token costs. Therefore, the text prices alone do not represent the complete cost of a voice application: audio exchanged during a session can materially affect billing.
The supplied editorial evaluation rates GPT-Realtime Mini highly for speed and cost, with scores of 8 and 9 respectively, while assigning lower scores of 5 for reasoning and 4 for coding. These are editorial assessments rather than OpenAI-published benchmark results or official capability ratings. They are best understood as a positioning signal: GPT-Realtime Mini is intended to favor responsive, affordable interaction over demanding reasoning or software-development workloads.
Main strengths and limitations
GPT-Realtime Mini’s strongest feature is the combination of realtime audio interaction and relatively low listed text-token pricing. It can support a natural voice interface while also accepting image context and invoking application functions. The 32,000-token context window and 4,096-token maximum output provide useful room for ordinary conversational workflows.
Its limitations are equally important:
- Scheduled removal: OpenAI has announced API shutdown for January 20, 2027.
- No video input: The model accepts images but not video.
- No structured outputs: Applications requiring provider-enforced schema formatting cannot rely on this feature.
- No fine-tuning: The model cannot be fine-tuned for a specialized domain through the supported offering.
- Audio costs are separate from the headline text prices: Realtime voice usage can involve modality-specific audio-token billing.
- Not optimized for the most demanding reasoning or coding: The supplied editorial scores position those capabilities below its speed and cost characteristics.
The research also identifies the model’s streaming feature flag as unsupported, even though the model supports realtime connections through WebRTC, WebSocket, and SIP. Developers should distinguish the documented realtime transport and session capabilities from any separate streaming feature label when checking compatibility.
When to choose GPT-Realtime Mini
GPT-Realtime Mini can be a reasonable choice when the immediate requirement is a lower-cost realtime voice experience and the integration will be migrated before the announced shutdown. It is particularly well suited to:
- Voice agents that answer questions or guide users through a process.
- Customer-support interfaces where spoken input and spoken replies are central.
- Speech-to-speech assistants that also need text transcripts or displayed responses.
- Realtime applications that call business functions during a conversation.
- Multimodal assistants that combine audio with occasional image input.
- Prototypes or existing deployments where speed and cost are more important than advanced reasoning.
It is less appropriate for a new long-lived production system because the model is deprecated. A team beginning a new integration should evaluate GPT-Realtime-2.1 Mini, which OpenAI identifies as the recommended replacement. The supplied research does not provide a full specification or price comparison for that replacement, so the newer model’s exact trade-offs should be verified separately rather than assumed.
When another option is more appropriate
Choose another option when the project requires video understanding, image generation, structured outputs, fine-tuning, or a stable lifecycle beyond January 2027. A model with stronger reasoning or coding performance may also be preferable for complex planning, software engineering, or tasks where response quality matters more than low latency and cost.
GPT-Realtime Mini can still make sense for an existing system that already depends on its realtime behavior, but that system should have a migration plan. The current alias is gpt-realtime-mini, which points to the gpt-realtime-mini-2025-12-15 snapshot. The earlier gpt-realtime-mini-2025-10-06 snapshot is deprecated. Pinning behavior, testing the recommended successor, and checking audio-token costs are practical steps before committing further resources.
Release history and current status
OpenAI released GPT-Realtime Mini on October 6, 2025. The current alias points to the December 15, 2025 snapshot. OpenAI marked the model family for deprecation on July 20, 2026 and announced API shutdown on January 20, 2027.
In practical terms, GPT-Realtime Mini is a capable low-cost realtime voice model with multimodal input and function calling, but it should be treated as a transitional or migration-bound choice rather than OpenAI’s preferred long-term starting point. Its value is clearest when an application needs responsive audio interaction now and can account for both audio billing and the announced end-of-service date.

