What is GPT-4o Mini Realtime?
GPT-4o Mini Realtime is OpenAI’s compact realtime model for applications that need low-latency conversations involving speech. Instead of treating audio as a separate transcription and synthesis pipeline, a realtime model can receive audio and return audio within an ongoing interaction. This makes it a practical fit for voice-driven software such as assistants, conversational customer support, interactive tutoring, and audio-first prototypes.
The model was released as a preview under the API identifier gpt-4o-mini-realtime-preview. Its dated snapshot was gpt-4o-mini-realtime-preview-2024-12-17. OpenAI now lists the GPT-4o Mini Realtime family as deprecated. The dated snapshot was separately deprecated and removed on May 7, 2026, while the canonical family is scheduled for API removal on January 20, 2027.
For new projects, OpenAI recommends GPT-Realtime-2.1 Mini instead. GPT-4o Mini Realtime remains relevant when evaluating an existing integration, understanding an older realtime deployment, or planning a migration from its API identifier.
Audio and text modalities
GPT-4o Mini Realtime supports both text and audio as inputs and outputs. In practical terms, an application can send written messages or speech and receive written replies or generated speech. This is different from a text-only language model, where an application would need separate speech-recognition and text-to-speech systems to create a voice conversation.
- Text input: Supported.
- Audio input: Supported.
- Text output: Supported.
- Audio output: Supported.
- Image input: Not supported.
- Video input: Not supported.
- Image or video generation: Not supported.
Realtime access is provided through WebRTC or WebSocket connections. WebRTC is commonly associated with browser and client-side realtime communication, while WebSocket connections are useful for persistent application-to-server sessions. The supplied model documentation identifies these interfaces as supported access methods; it does not describe ordinary streaming as a separate capability, because realtime interaction is the model’s central operating pattern.
Context window and output limits
The model has a 16,000-token context window and a maximum output of 4,096 tokens. A token is a piece of text used internally by the model; the context window is the total amount of conversation and other input that can be considered in one request. Audio interactions also consume model tokens according to OpenAI’s audio-token accounting.
| Specification | GPT-4o Mini Realtime |
|---|---|
| Provider | OpenAI |
| Model identifier | gpt-4o-mini-realtime-preview |
| Release snapshot | December 17, 2024 |
| Context window | 16,000 tokens |
| Maximum output | 4,096 tokens |
| Knowledge cutoff | October 1, 2023 |
| Function calling | Supported |
| Structured outputs | Not supported |
| Fine-tuning | Not supported |
| Web search | Not documented as supported |
The October 1, 2023 knowledge cutoff means the model should not be expected to know later events or current information on its own. An application that needs up-to-date answers must provide relevant information through its own context or use another supported retrieval system. Realtime connectivity does not update the model’s underlying knowledge.
GPT-4o Mini Realtime pricing
OpenAI prices GPT-4o Mini Realtime according to both text and audio token usage. Audio tokens are substantially more expensive than text tokens, so a voice-heavy application should estimate costs using its expected speaking and listening volume rather than relying only on text-token rates.
| Usage type | Price per 1 million tokens |
|---|---|
| Text input | $0.60 |
| Text output | $2.40 |
| Audio input | $10.00 |
| Audio output | $20.00 |
| Cached input | $0.30 for both text and audio, as listed |
These are usage prices rather than a fixed subscription fee. The difference between text and audio pricing is important: an application that primarily exchanges short text messages may cost much less than one that continuously processes microphone input and generates spoken responses. Caching can reduce the price of eligible repeated input, but the supplied documentation does not establish that every part of a realtime conversation will qualify for cached pricing.
Where the model is strongest
GPT-4o Mini Realtime’s main advantage is the combination of realtime audio support, relatively low pricing, and a lightweight interaction profile. It is designed for conversations where response speed and a natural turn-taking experience matter more than advanced reasoning or broad multimodal understanding.
- Voice assistants: The model can accept spoken requests and return spoken answers without requiring a text-only interaction design.
- Interactive audio products: Web and mobile applications can use its realtime interfaces for voice-driven workflows.
- Customer support prototypes: Function calling can let the model request application actions, such as looking up information or initiating a supported workflow.
- Tutoring and guided conversations: Audio input and output are useful when users need to speak naturally rather than type.
- Cost-sensitive experiments: The model’s lower text and audio prices make it suitable for testing a voice experience before moving to a more capable or newer model.
The research characterizes it as fast and inexpensive relative to more capable options. The comparative editorial scores supplied for the model are 9 out of 10 for speed and 8 out of 10 for cost. These are editorial estimates, not OpenAI-published benchmark results, and should be treated as directional rather than as measured guarantees for a particular application.
Reasoning, coding, and function support
GPT-4o Mini Realtime is a specialized realtime audio model, not a frontier reasoning model. It can follow conversational instructions and support application interactions, but the supplied research does not provide a benchmark demonstrating advanced mathematical, analytical, or long-horizon reasoning performance.
The same distinction applies to coding. The editorial coding score is 3 out of 10, while the reasoning score is 4 out of 10. These scores are subjective comparative evaluations rather than vendor-published test results. They suggest that the model should be selected for conversational audio tasks first, not as the primary engine for complex software development or difficult technical analysis.
Function calling is supported. This allows an application to expose defined functions or tools that the model can request during a conversation. For example, a voice assistant might ask an application to retrieve an account detail or start a supported business process. Function calling does not mean that the model can freely execute arbitrary operations: the surrounding application must define the available functions, validate arguments, enforce permissions, and perform the actual action.
Important limitations
The model’s limitations affect both capability selection and product planning:
- No visual input: It cannot directly handle image or video workflows according to the supplied specifications.
- No structured outputs: Applications that require provider-enforced JSON or another strict output schema should choose an option that explicitly supports that feature.
- No fine-tuning: The supplied documentation does not list fine-tuning support for this model.
- No documented web search: It should not be treated as having built-in access to current web information.
- Old knowledge cutoff: Its knowledge cutoff is October 1, 2023, so current facts must come from application-provided context or another retrieval method.
- Moderate context capacity: A 16,000-token context window may be sufficient for many conversations but can be restrictive for large documents, long-running sessions, or extensive tool history.
- Deprecation: New production systems face a planned migration because the family is scheduled for API removal on January 20, 2027.
These limitations are especially significant for products that combine voice with visual understanding, require machine-readable structured responses, or depend on current information. In those cases, a newer realtime model or a separate retrieval and processing architecture may be more appropriate.
When to choose GPT-4o Mini Realtime
Choose GPT-4o Mini Realtime when you are maintaining an existing integration or evaluating a low-cost, speech-first interaction and can accommodate its deprecation timeline. It is most defensible for prototypes, internal experiments, short-lived demonstrations, and legacy applications where text-and-audio conversation is more important than image understanding, advanced reasoning, or strict structured output.
Its cost and speed profile can also make sense when the main product requirement is responsive voice interaction at modest complexity. For example, a simple spoken assistant that answers from application-supplied information and invokes a small number of controlled functions may not need a larger general-purpose model.
Do not choose it for a new long-lived production deployment unless there is a specific migration or compatibility reason. OpenAI recommends GPT-Realtime-2.1 Mini as its replacement. A newer model is also preferable when the application needs a supported future lifecycle, more capable reasoning, current information workflows, visual input, structured output, or a broader range of multimodal behavior. The exact capabilities and pricing of the replacement are not established by the supplied research, so they should be verified separately before migration.
Availability and migration considerations
The canonical alias is gpt-4o-mini-realtime-preview. The dated identifier gpt-4o-mini-realtime-preview-2024-12-17 should not be treated as an available current target because the supplied lifecycle information states that it was removed on May 7, 2026.
Teams using the canonical family should inventory the parts of their application that depend on its realtime session behavior, audio input and output handling, function definitions, token accounting, and error handling. They should then test the recommended GPT-Realtime-2.1 Mini replacement with representative conversations, interruptions, tool calls, audio conditions, and expected costs. Since migration details are not provided here, compatibility should be confirmed against OpenAI’s current documentation rather than assumed from the shared model name.
GPT-4o Mini Realtime is therefore best understood as a historical and still-transitioning low-cost realtime model: useful for understanding or operating an existing voice application, but not the safest default for a new system expected to run beyond its announced shutdown date.

