GPT-Realtime-2.1

GPT-Realtime-2.1 Mini

by OpenAI · Current

OpenAI's GPT-Realtime-2.1 Mini is a lower-cost real-time reasoning model for voice agents. It supports text, audio, and image input, text and audio output, function calling, and WebRTC, WebSocket, and SIP connections. The model has a 128,000-token context window, a 32,000-token maximum output, and a September 30, 2024 knowledge cutoff. It does not support video input, structured outputs, or fine-tuning.

Text Speech Reasoning Coding
GPT-Realtime-2.1 Mini is designed for developers who need responsive, speech-capable AI agents without the cost of OpenAI's full-size GPT-Realtime-2.1 model. It supports live text and audio interaction, image input, and tool calls, making it suitable for customer-service agents, tutors, scheduling assistants, and other applications that need natural conversations with access to external systems.
Outputs

What GPT-Realtime-2.1 Mini can produce

Text Speech
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Tool use Prompt caching Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime-2.1
Model type Lightweight
Context window 128K tokens
Maximum output 32K tokens
Knowledge cutoff September 30, 2024
Release date 2026-07-06
Status Current
Knowledge cutoff notes

The official model documentation lists September 30, 2024 as the knowledge cutoff. Realtime connections, tool calls, and application-provided context do not change the underlying cutoff.

Model notes

GPT-Realtime-2.1 Mini is a distilled reasoning model optimized for faster and lower-cost realtime voice interactions. It supports text and audio input/output, image input, function calling, and connections through WebRTC, WebSocket, and SIP. Video input, structured outputs, fine-tuning, and image or video generation are not supported. The documented knowledge cutoff is September 30, 2024. Editorial scores are comparative estimates rather than provider-published benchmarks.

Cost

Model pricing

Input Text: $0.60 per 1M tokens; cached text: $0.06 per 1M; audio: $10.00 per 1M tokens; cached audio: $0.30 per 1M; image: $0.80 per 1M tokens; cached image: $0.08 per 1M
Output Text: $2.40 per 1M tokens; audio: $20.00 per 1M tokens
Model guide

GPT-Realtime-2.1 Mini: Features, Pricing, Limits and API Capabilities

GPT-Realtime-2.1 Mini is OpenAI's faster, lower-cost distilled reasoning model for real-time voice applications. It accepts text, audio, and images, produces text and audio, supports function calling, and connects through WebRTC, WebSocket, and SIP. It has a 128,000-token context window and a 32,000-token maximum output, but does not support video input, structured outputs, fine-tuning, or standard streaming.

What is GPT-Realtime-2.1 Mini?

GPT-Realtime-2.1 Mini is an OpenAI model built for real-time conversational applications. It is a distilled reasoning model: compared with a larger model in the same family, it is positioned for faster responses and lower operating costs rather than maximum general-purpose capability. Its main focus is live voice interaction, including speech-to-speech conversations in which an application sends audio and receives spoken audio in response.

The model is listed as a current member of OpenAI's GPT-Realtime-2.1 family. The provider's documentation describes it as a lightweight real-time model, while the comparative scores supplied for this page are editorial estimates rather than OpenAI-published benchmarks. Those estimates rate its reasoning at 7 out of 10, coding at 5 out of 10, speed at 9 out of 10, and cost efficiency at 8 out of 10. They should be treated as directional guidance, not standardized performance measurements.

GPT-Realtime-2.1 Mini is not simply a text model with a separate speech-to-text layer. Audio input and audio output are supported as part of its real-time use case, allowing developers to build conversational agents that listen and speak within a live session.

Capabilities and supported modalities

The model accepts text, audio, and images as input. It can return text and audio as output. This combination supports both conventional text interaction and speech-to-speech applications, while image input allows an agent to use visual information supplied by the user or application.

  • Text input and output: Supported.
  • Audio input and output: Supported for real-time voice interactions.
  • Image input: Supported for visual context.
  • Video input: Not supported.
  • Image, video, music, and embedding output: Not supported.
  • Function calling: Supported.
  • Structured outputs: Not supported according to the supplied model documentation.

In practical terms, a voice assistant could receive spoken questions, consult an external scheduling or customer-record system through a function call, and answer with synthesized audio. An application could also provide an image as context, but GPT-Realtime-2.1 Mini should not be selected for workflows that require it to understand a video stream or generate images or video.

Context window and output limit

GPT-Realtime-2.1 Mini has a 128,000-token context window and a maximum output length of 32,000 tokens. A context window is the amount of conversation and other text the model can consider in a request or session. It can include earlier turns, instructions, tool results, and other application-provided information, subject to the endpoint and session behavior used by the developer.

The large context window is useful for longer-running voice sessions, support histories, reference material, or tool output that must remain available during a conversation. It does not mean that every request should include the entire history. Sending unnecessary context increases processing requirements and can raise costs, especially when audio is involved.

The documented knowledge cutoff is September 30, 2024. A real-time connection, tool call, or application-provided document can give the agent current information, but those mechanisms do not change the model's underlying knowledge cutoff.

GPT-Realtime-2.1 Mini pricing

Pricing is usage-based and varies by modality. The supplied OpenAI pricing information lists the following rates per one million tokens:

Usage typePrice per 1M tokens
Text input$0.60
Cached text input$0.06
Text output$2.40
Audio input$10.00
Cached audio input$0.30
Audio output$20.00
Image input$0.80
Cached image input$0.08

There is no recurring subscription price for the model in the supplied research. These are API usage rates, and the amount a project pays depends on the number of input and output tokens consumed. Audio is substantially more expensive than text: audio input costs more than sixteen times the price of text input, while audio output costs more than eight times the price of text output. Developers planning a voice application should therefore estimate speech volume, conversation length, repeated context, interruption frequency, and output duration instead of using text-only pricing as a proxy.

Cached input pricing can reduce the cost of repeatedly supplied context when the applicable API behavior allows that context to be reused. It does not make all audio or text usage automatically inexpensive, and the exact billing behavior should be checked against the endpoint being used.

Real-time API connections and integrations

GPT-Realtime-2.1 Mini is intended for OpenAI's real-time interfaces. The supplied documentation identifies WebRTC, WebSocket, and SIP connection options:

  • WebRTC: A suitable connection style for browser and client applications where interactive audio is important.
  • WebSocket: Useful for server-side integrations that need a persistent real-time connection.
  • SIP: Relevant to telephony and phone-based conversational workflows.

Function calling lets the model request an operation from the surrounding application rather than directly changing a database or service on its own. For example, a voice agent could ask the application to check appointment availability, retrieve an order status, or create a reservation. The application remains responsible for validating the request, applying permissions, performing the operation, and returning the result to the model.

The model is listed across real-time and related text and audio API surfaces, but behavior can differ by endpoint. Developers should confirm which modalities, connection methods, event types, and tool features are available for the specific integration they choose. The supplied research also states that standard streaming is not supported as a model capability, so real-time connection support should not be assumed to mean that every conventional streaming interface is available.

Reasoning, coding, speed and cost trade-offs

OpenAI positions GPT-Realtime-2.1 Mini as a faster and less expensive alternative to the full-size GPT-Realtime-2.1. The trade-off is that the Mini model is not intended to provide the strongest reasoning performance in the family. It is better suited to handling routine conversational decisions, classification, retrieval through tools, and short multi-step interactions than to difficult analysis requiring the highest available reasoning quality.

The supplied editorial assessment gives the model a high speed score and a moderate coding score. The coding score is not a provider benchmark and should not be read as a guarantee of software-engineering performance. GPT-Realtime-2.1 Mini can use tools and generate text, but it is not described as a dedicated coding model. For complex code generation, large refactoring tasks, or demanding software-engineering workflows, a specialized or stronger general-purpose model may be more appropriate.

The principal cost-versus-capability decision is modality-dependent. A text-heavy assistant can be relatively inexpensive at the listed text rates. A voice-heavy assistant pays much more for audio tokens, even though the Mini model is cheaper than the full-size real-time option. Lower model cost therefore does not remove the need to control conversation length, reuse context where appropriate, and avoid unnecessary audio processing.

Main strengths

  • Low-latency orientation: The model is designed for live interactions where waiting for a long response harms the user experience.
  • Native speech interaction: Audio input and output support speech-to-speech assistants rather than requiring a text-only interface.
  • Lower cost than the full-size real-time model: This makes it more practical for high-volume or cost-sensitive voice deployments, although audio remains relatively expensive.
  • Tool use: Function calling can connect conversations to business systems, scheduling services, databases, and other application functions.
  • Image context: Image input extends a voice or text conversation beyond language alone.
  • Large context: The 128,000-token window can accommodate substantial conversation history and application context.

Important limitations

  • It does not support video input.
  • It does not generate images, video, music, or embeddings.
  • Structured outputs are not supported in the supplied model documentation, so it should not be chosen when guaranteed JSON-schema responses are a core requirement.
  • Fine-tuning is not supported.
  • It is not positioned as OpenAI's strongest reasoning model or as a dedicated coding model.
  • Audio input and output cost substantially more than text input and output.
  • The model's knowledge cutoff is September 30, 2024; current information must come from tools or application-provided context.
  • Connection and feature behavior may vary between real-time and related API surfaces, so endpoint-specific documentation remains important.

Best use cases

GPT-Realtime-2.1 Mini is a good fit when the application needs a responsive conversational voice experience and does not require the maximum reasoning capability available. Suitable examples include customer-support triage, appointment and reservation agents, language-learning tutors, voice-controlled business tools, interactive help desks, and assistants that retrieve information or perform approved actions through functions.

It is especially appropriate when response speed and operating cost matter more than solving the hardest possible reasoning problems. A business handling many short voice conversations may prefer the Mini model because its lower cost and real-time focus can matter more than a larger model's additional reasoning capacity.

When to choose GPT-Realtime-2.1 Mini

Choose GPT-Realtime-2.1 Mini when you need all or most of the following:

  • Live audio input and audio output.
  • Fast conversational responses.
  • Function calling to external systems.
  • Optional image context.
  • A lower-cost alternative to a larger real-time reasoning model.
  • A context window large enough for extended conversations or substantial application context.

Choose another option when the primary requirement is maximum reasoning quality, advanced coding, fine-tuning, video understanding, image or video generation, or guaranteed structured JSON responses. The full-size GPT-Realtime-2.1 may be a better comparison when quality is more important than cost and latency. A text-focused model may be more economical for applications that do not need speech, while a model with structured-output support is more suitable for workflows that depend on strict machine-readable responses.

Overall, GPT-Realtime-2.1 Mini is best understood as a practical real-time voice model rather than a universal replacement for every OpenAI model. Its value comes from combining responsive speech interaction, image input, tool use, and a relatively lower operating cost in one model. The main compromises are weaker positioning for demanding reasoning and coding tasks, the absence of structured outputs and fine-tuning, and the higher price of audio processing.


Answers to Frequently Asked Questions

What are the main limitations of GPT-Realtime-2.1 Mini?
The model does not support video input, image or video generation, structured outputs, or fine-tuning. It is not positioned as OpenAI's strongest reasoning or coding model, and its knowledge cutoff is September 30, 2024. Endpoint-specific support for modalities and features should also be verified in the relevant API documentation.
What APIs and connection methods work with GPT-Realtime-2.1 Mini?
GPT-Realtime-2.1 Mini is intended for OpenAI real-time interfaces using WebRTC, WebSocket, or SIP. WebRTC suits browser and client applications, WebSocket is useful for server-side integrations, and SIP supports telephony workflows. The model also supports function calling.
How much does GPT-Realtime-2.1 Mini cost?
The listed API prices per one million tokens are $0.60 for text input, $0.06 for cached text input, $2.40 for text output, $10.00 for audio input, $0.30 for cached audio input, $20.00 for audio output, $0.80 for image input, and $0.08 for cached image input. Costs depend on actual token usage, and audio is significantly more expensive than text.
What is GPT-Realtime-2.1 Mini designed for?
GPT-Realtime-2.1 Mini is a lightweight OpenAI model for real-time conversational applications, especially live voice assistants. It supports speech-to-speech interaction, fast responses, function calling, and lower operating costs than the full-size GPT-Realtime-2.1.
What input and output modalities does GPT-Realtime-2.1 Mini support?
The model accepts text, audio, and images as input and can return text and audio as output. It does not support video input or generate images, video, music, or embeddings.


Sources 5
Provider

About OpenAI