GPT-Realtime

GPT-Realtime Mini

by OpenAI · Deprecated; currently accessible but scheduled for API shutdown on 2027-01-20

OpenAI’s lower-cost realtime model for voice agents and multimodal assistants. It accepts text, images, and audio, returns text and audio, supports function calling, and has a 32,000-token context window. It is deprecated and scheduled for API shutdown on January 20, 2027.

Text Speech Reasoning Coding
GPT-Realtime Mini is designed for applications that need responsive speech-to-speech interaction without the cost of OpenAI’s larger realtime options. It can handle text, image, and audio input, produce text and audio responses, and connect through WebRTC, WebSocket, or SIP. Its combination of low listed text-token pricing, audio support, and function calling makes it suitable for voice agents and interactive assistants. However, its scheduled shutdown makes lifecycle planning essential.
Outputs

What GPT-Realtime Mini can produce

Text Speech
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Tool use Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

5/10 Reasoning
4/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime
Model type Realtime
Context window 32K tokens
Maximum output 4K tokens
Knowledge cutoff 2023-10-01
Release date 2025-10-06
Status Deprecated; currently accessible but scheduled for API shutdown on 2027-01-20
Deprecation date 2026-07-20
Shutdown date 2027-01-20
Knowledge cutoff notes

The official model page lists October 1, 2023 as the knowledge cutoff. Realtime access, external tools, or supplied context do not change the underlying cutoff.

Model notes

Released on October 6, 2025. The current gpt-realtime-mini alias points to gpt-realtime-mini-2025-12-15; the 2025-10-06 snapshot is deprecated. OpenAI announced deprecation of the model family on July 20, 2026 and API shutdown on January 20, 2027, with GPT-Realtime-2.1 Mini as the recommended replacement. The model accepts text, images, and audio, and returns text and audio. The model page lists function calling as supported, but structured outputs, fine-tuning, and the streaming feature flag as unsupported. Realtime sessions can incur modality-specific audio token costs in addition to the listed text-token prices.

Cost

Model pricing

Input $0.60 per 1M text input tokens; $0.06 per 1M cached text input tokens
Output $2.40 per 1M text output tokens
Model guide

GPT-Realtime Mini: Lower-Cost Voice AI with a Scheduled Shutdown

GPT-Realtime Mini is OpenAI’s lower-cost realtime model for interactive voice and multimodal applications. It accepts text, images, and audio, returns text and audio, supports function calling, and provides a 32,000-token context window. Although it remains accessible, it is deprecated and scheduled for API shutdown on January 20, 2027, so new projects should evaluate GPT-Realtime-2.1 Mini instead.

What is GPT-Realtime Mini?

GPT-Realtime Mini is an OpenAI realtime model for interactive applications, particularly those that communicate through voice. Unlike a conventional text-only model interaction, a realtime session can receive audio and return spoken audio with low conversational delay. It can also work with text and image inputs, allowing an application to combine spoken conversation with written instructions or visual context.

The model is positioned as a smaller, more cost-efficient member of OpenAI’s realtime model family. Its main appeal is not maximum reasoning performance; it is the balance between responsiveness, multimodal input, audio output, and operating cost. Typical applications include voice agents, customer-service interfaces, speech-to-speech assistants, and realtime interfaces that use function calling to access application actions or data.

There is an important lifecycle qualification: GPT-Realtime Mini is deprecated. OpenAI has announced API shutdown for January 20, 2027, and recommends GPT-Realtime-2.1 Mini as the replacement. It may still be useful for existing integrations or short-term evaluation, but the planned shutdown should be treated as a central factor in any production decision.

Supported input and output modalities

GPT-Realtime Mini accepts three input types: text, images, and audio. It produces text and audio output. This makes it suitable for both speech-to-speech conversations and applications where users speak but the interface displays a written answer as well.

  • Text input: Supported.
  • Image input: Supported for visual context, but the model does not generate images.
  • Audio input: Supported for spoken interaction and realtime voice sessions.
  • Video input: Not supported.
  • Text output: Supported.
  • Audio output: Supported, including spoken responses.
  • Image and video output: Not supported.

For example, a voice assistant could receive a spoken question, inspect an image supplied by the user, call a business function, and return both a spoken response and text for the application interface. The supplied research does not indicate that GPT-Realtime Mini can directly process video streams, so video-based use cases require another design or model.

Context window and technical specifications

The model has a 32,000-token context window. A context window is the amount of conversation, instructions, tool information, and other model-visible content that can be considered within a session. This is large enough for many ordinary voice-agent interactions, although long-running sessions and applications that attach substantial documents or tool output will need to manage context carefully.

SpecificationGPT-Realtime Mini
ProviderOpenAI
Model familyGPT-Realtime
Model typeRealtime
Context window32,000 tokens
Maximum output4,096 tokens
InputText, image, and audio
OutputText and audio
Function callingSupported
Fine-tuningNot supported
Structured outputsNot supported

The maximum output limit is 4,096 tokens. In a voice application, this limit applies to the model response rather than directly describing the duration of an audio clip; actual audio usage and response length depend on the session and modality. The model supports function calling, which allows an application to expose defined operations such as checking an order, booking an appointment, or retrieving account information. The model can request a function, but the surrounding application remains responsible for executing it and handling the result.

Realtime connections and API support

GPT-Realtime Mini is documented for OpenAI’s Realtime API and related session workflows. Connections can use WebRTC, WebSocket, or SIP. WebRTC is commonly associated with browser or client voice experiences, WebSocket with server-managed connections, and SIP with telephony integrations; the appropriate choice depends on the architecture of the application.

The model documentation lists support across several OpenAI endpoint categories, including Chat Completions, Responses, Live sessions, realtime translation, realtime transcription sessions, and Batch. The model should still be evaluated in the specific endpoint and session pattern an application needs, because a model being listed for an endpoint does not remove the practical differences between ordinary request-response use and a realtime audio session.

GPT-Realtime Mini supports function calling but not structured outputs. Structured outputs are useful when an application requires the model to conform to a formally defined response schema. Since that capability is not supported here, developers who need reliable schema-constrained responses may need additional validation, a different workflow, or another model.

Pricing and cost trade-offs

OpenAI’s listed text-token prices are:

  • Text input: $0.60 per 1 million tokens.
  • Cached text input: $0.06 per 1 million tokens.
  • Text output: $2.40 per 1 million tokens.

These figures make GPT-Realtime Mini a cost-conscious option for text-token usage. Cached input is substantially cheaper than uncached input, which may matter when applications repeatedly send stable instructions or other reusable context. The supplied pricing information also warns that realtime sessions can incur modality-specific audio-token costs. Therefore, the text prices alone do not represent the complete cost of a voice application: audio exchanged during a session can materially affect billing.

The supplied editorial evaluation rates GPT-Realtime Mini highly for speed and cost, with scores of 8 and 9 respectively, while assigning lower scores of 5 for reasoning and 4 for coding. These are editorial assessments rather than OpenAI-published benchmark results or official capability ratings. They are best understood as a positioning signal: GPT-Realtime Mini is intended to favor responsive, affordable interaction over demanding reasoning or software-development workloads.

Main strengths and limitations

GPT-Realtime Mini’s strongest feature is the combination of realtime audio interaction and relatively low listed text-token pricing. It can support a natural voice interface while also accepting image context and invoking application functions. The 32,000-token context window and 4,096-token maximum output provide useful room for ordinary conversational workflows.

Its limitations are equally important:

  • Scheduled removal: OpenAI has announced API shutdown for January 20, 2027.
  • No video input: The model accepts images but not video.
  • No structured outputs: Applications requiring provider-enforced schema formatting cannot rely on this feature.
  • No fine-tuning: The model cannot be fine-tuned for a specialized domain through the supported offering.
  • Audio costs are separate from the headline text prices: Realtime voice usage can involve modality-specific audio-token billing.
  • Not optimized for the most demanding reasoning or coding: The supplied editorial scores position those capabilities below its speed and cost characteristics.

The research also identifies the model’s streaming feature flag as unsupported, even though the model supports realtime connections through WebRTC, WebSocket, and SIP. Developers should distinguish the documented realtime transport and session capabilities from any separate streaming feature label when checking compatibility.

When to choose GPT-Realtime Mini

GPT-Realtime Mini can be a reasonable choice when the immediate requirement is a lower-cost realtime voice experience and the integration will be migrated before the announced shutdown. It is particularly well suited to:

  • Voice agents that answer questions or guide users through a process.
  • Customer-support interfaces where spoken input and spoken replies are central.
  • Speech-to-speech assistants that also need text transcripts or displayed responses.
  • Realtime applications that call business functions during a conversation.
  • Multimodal assistants that combine audio with occasional image input.
  • Prototypes or existing deployments where speed and cost are more important than advanced reasoning.

It is less appropriate for a new long-lived production system because the model is deprecated. A team beginning a new integration should evaluate GPT-Realtime-2.1 Mini, which OpenAI identifies as the recommended replacement. The supplied research does not provide a full specification or price comparison for that replacement, so the newer model’s exact trade-offs should be verified separately rather than assumed.

When another option is more appropriate

Choose another option when the project requires video understanding, image generation, structured outputs, fine-tuning, or a stable lifecycle beyond January 2027. A model with stronger reasoning or coding performance may also be preferable for complex planning, software engineering, or tasks where response quality matters more than low latency and cost.

GPT-Realtime Mini can still make sense for an existing system that already depends on its realtime behavior, but that system should have a migration plan. The current alias is gpt-realtime-mini, which points to the gpt-realtime-mini-2025-12-15 snapshot. The earlier gpt-realtime-mini-2025-10-06 snapshot is deprecated. Pinning behavior, testing the recommended successor, and checking audio-token costs are practical steps before committing further resources.

Release history and current status

OpenAI released GPT-Realtime Mini on October 6, 2025. The current alias points to the December 15, 2025 snapshot. OpenAI marked the model family for deprecation on July 20, 2026 and announced API shutdown on January 20, 2027.

In practical terms, GPT-Realtime Mini is a capable low-cost realtime voice model with multimodal input and function calling, but it should be treated as a transitional or migration-bound choice rather than OpenAI’s preferred long-term starting point. Its value is clearest when an application needs responsive audio interaction now and can account for both audio billing and the announced end-of-service date.


Answers to Frequently Asked Questions

What are the main limitations of GPT-Realtime Mini?
GPT-Realtime Mini does not support video input, structured outputs, or fine-tuning. It is also less suitable for demanding reasoning and coding tasks, and its realtime audio usage can add costs beyond the listed text-token prices.
Is GPT-Realtime Mini still supported, and when will it shut down?
GPT-Realtime Mini is deprecated, and OpenAI has announced API shutdown for January 20, 2027. OpenAI recommends GPT-Realtime-2.1 Mini as the replacement, so new long-term integrations should evaluate that model instead.
How much does GPT-Realtime Mini cost?
The listed prices are $0.60 per 1 million text input tokens, $0.06 per 1 million cached text input tokens, and $2.40 per 1 million text output tokens. Realtime voice applications may also incur separate audio-token charges, so text pricing does not represent the full cost of an audio session.
What input and output modalities does GPT-Realtime Mini support?
The model supports text, image, and audio input, and produces text and audio output. It does not support video input or image and video generation.
What is GPT-Realtime Mini used for?
GPT-Realtime Mini is designed for low-latency interactive applications such as voice agents, customer-service interfaces, speech-to-speech assistants, and realtime applications that use function calling. It accepts text, image, and audio input and can return text and audio output.


Sources 4
Provider

About OpenAI