What is GPT-Realtime-2.1?
GPT-Realtime-2.1 is OpenAI's current realtime reasoning model for interactive voice and multimodal applications. Unlike a conventional text model that receives a prompt and returns a completion, it is designed to maintain a live conversation in which audio can be exchanged continuously and the user can interrupt the assistant.
The model can understand text, speech, and images. It can respond with text or synthesized audio, allowing developers to build speech-to-speech experiences without treating transcription and speech generation as completely separate application steps. Function calling lets it invoke external business logic, such as looking up an account, checking an order, scheduling an appointment, or updating a customer record.
OpenAI released GPT-Realtime-2.1 on July 6, 2026, as an update to GPT-Realtime-2. Its primary role in OpenAI's model lineup is realtime reasoning for voice agents, rather than inexpensive bulk text generation, image creation, video understanding, or general-purpose embeddings.
What changed from GPT-Realtime-2?
According to the supplied OpenAI documentation, GPT-Realtime-2.1 improves several parts of voice interaction:
- Alphanumeric recognition: better handling of order numbers, account identifiers, confirmation codes, phone numbers, addresses, and similar information spoken aloud.
- Silence and background-noise handling: improved behavior when a caller pauses, speaks in a noisy environment, or takes time to respond.
- Interruption behavior: more natural handling of users who begin speaking while the assistant is still responding.
- Configurable reasoning effort: developers can adjust how much reasoning the model applies. Higher effort may help with difficult decisions and tool-use flows, but can increase latency and output-token consumption.
These improvements matter most in production voice applications, where a misheard identifier or poorly handled interruption can be more damaging than a slightly imperfect response in a casual conversation.
Capabilities and supported modalities
GPT-Realtime-2.1 supports multiple input and output types, but it is not a model for every kind of multimodal generation. Its supported modalities are:
| Capability | Support |
|---|---|
| Text input | Yes |
| Audio input | Yes |
| Image input | Yes |
| Text output | Yes |
| Audio output | Yes |
| Video input | No |
| Image generation | No |
| Video generation | No |
Audio input and output enable direct speech-to-speech conversations. Image input can be useful when a voice agent needs to discuss a document, photograph, diagram, or other visual supplied by the user. The model does not natively create images or video, and it does not accept video input.
Context window and output limits
GPT-Realtime-2.1 has a documented context window of 128,000 tokens and a maximum output of 32,000 tokens. The context window is the amount of conversation and other input the model can consider in a request or realtime session. In a long-running voice application, developers still need to manage conversation history because retaining every turn increases input processing and cost.
The maximum output is substantially larger than most spoken responses require. In practice, voice applications generally need short, timely turns rather than tens of thousands of generated tokens. The limit is more relevant to complex tool workflows, detailed text responses, or applications that use the model for both conversational and written output.
Reasoning, coding, and tool use
GPT-Realtime-2.1 supports configurable reasoning effort. This gives developers a choice between faster responses for routine dialogue and more deliberate processing for difficult decisions, multi-step instructions, or tool calls. The trade-off is not simply quality versus speed: higher reasoning effort can also increase output-token usage and therefore operating cost.
Function calling is supported. In a voice agent, a function is an application-defined operation that the model can request, while the application performs the actual task and returns the result. Examples include retrieving a customer record, checking inventory, calculating a quote, or creating a calendar appointment. The model can then explain the result in the ongoing conversation.
The supplied research rates its coding capability as a moderate-to-strong editorial assessment rather than a provider-published benchmark. GPT-Realtime-2.1 can participate in coding-related tool workflows and generate text, but it is primarily optimized for realtime voice interaction. Teams focused on large codebases, long coding sessions, or text-only software development may prefer a model designed specifically for those workloads.
How the Realtime API fits in
GPT-Realtime-2.1 is used with OpenAI's Realtime API for stateful conversational sessions. The model identifier is gpt-realtime-2.1. Realtime sessions can use WebRTC or WebSocket transports, with the application managing session configuration, audio exchange, events, interruptions, and tool results.
WebRTC is commonly associated with interactive browser or client audio connections, while WebSocket connections are useful for server-managed realtime communication. The exact transport choice depends on the application architecture. In either case, the model's realtime behavior is based on streamed events and audio exchange rather than a single request followed by one final response.
Developers should distinguish realtime event streaming from the model's listed feature support. The supplied model information says standard streaming is not listed as a separate supported feature, while Realtime API sessions themselves exchange streamed events and audio. GPT-Realtime-2.1 also does not support structured outputs, so applications that require guaranteed schema-conforming JSON should consider another model or add their own validation and recovery layer.
GPT-Realtime-2.1 pricing
GPT-Realtime-2.1 uses usage-based pricing that varies by modality. The supplied OpenAI pricing information lists the following rates per 1 million units:
| Usage type | Price |
|---|---|
| Text input | $4.00 per 1 million tokens |
| Cached text input | $0.40 per 1 million tokens |
| Text output | $24.00 per 1 million tokens |
| Audio input | $32.00 per 1 million audio tokens |
| Cached audio input | $0.40 per 1 million audio tokens |
| Audio output | $64.00 per 1 million audio tokens |
| Image input | $5.00 per 1 million tokens |
| Cached image input | $0.50 per 1 million tokens |
Audio output is the most expensive listed output modality, and audio input is also much more expensive than ordinary text input. This makes the model a better fit for applications where realtime speech is central to the product than for high-volume text processing. Caching repeated context and limiting unnecessary conversation history can reduce recurring input costs.
Main strengths and trade-offs
GPT-Realtime-2.1's main strength is the combination of realtime audio, reasoning, multimodal input, and external tool use. A customer-service agent can listen to a caller, identify a spoken account number, consult a business system, and respond naturally without forcing the user through a separate transcription or menu-driven workflow.
Its improvements to silence, noise, alphanumeric recognition, and interruption behavior are particularly useful in telephony and contact-center environments. The model is also more suitable than a basic speech interface when the conversation requires decisions, policy interpretation, or several tool calls.
The principal trade-off is cost and complexity. Audio processing costs more than text processing, and realtime applications must handle transport connections, session state, interruptions, audio events, and tool results. Higher reasoning effort can improve difficult workflows but may make responses slower and more expensive. These costs should be evaluated against the value of immediate spoken interaction.
Best use cases
- Customer-service and contact-center agents: handle spoken requests, retrieve records, and provide responses while maintaining a natural conversation.
- Telephony and SIP-based systems: support callers who may speak over the assistant, pause frequently, or provide numbers and codes in noisy conditions.
- Voice-operated business applications: connect speech interaction to scheduling, order management, account systems, or internal workflows.
- Realtime support assistants: combine conversation with external information retrieval and action-taking tools.
- Multimodal voice assistants: let users discuss an image, document, or other visual while speaking with the assistant.
- Complex conversational workflows: use configurable reasoning when the assistant must interpret instructions before selecting a tool or response.
Limitations to consider
GPT-Realtime-2.1 does not support video input, image generation, video generation, structured outputs, fine-tuning, native embeddings, or web search as a listed built-in capability. It is therefore not a complete choice for applications that need visual video analysis, guaranteed JSON responses, custom model training, or vector embedding generation.
The documented knowledge cutoff is September 30, 2024. Connecting the model to external tools can provide newer information, but tool access does not change the model's underlying training cutoff. Applications that answer questions about current events, prices, inventory, policies, or account data should use an appropriate live data source and should not rely on model memory alone.
The model's audio pricing also makes it less attractive for inexpensive, high-volume, text-only workloads. A text-focused model may be a better choice when low cost and throughput matter more than natural voice interaction.
When to choose GPT-Realtime-2.1
Choose GPT-Realtime-2.1 when the product needs a live voice conversation and the assistant must do more than transcribe speech or read back scripted responses. It is especially appropriate when the system must recognize precise spoken information, tolerate interruptions, reason about a request, and call external tools during the same interaction.
Consider another type of model when the task is primarily text generation, inexpensive classification, embeddings, structured data production, image generation, or video understanding. A specialized text or coding model may offer a better speed-and-cost profile for those workloads. Likewise, if the application only needs speech recognition and speech synthesis with minimal reasoning, a simpler pipeline may be easier to operate and less expensive.
For teams comparing OpenAI model options, [GPT-4o](/openai/models/gpt-4o) can provide useful context for a general multimodal model, while [GPT-5.4](/openai/models/gpt-5-4) is a more relevant comparison for applications prioritizing general reasoning and text-based capability rather than realtime speech. These alternatives should not be treated as drop-in replacements: GPT-Realtime-2.1's distinguishing value is its realtime voice-agent workflow.
Bottom line
GPT-Realtime-2.1 is designed for production voice agents that need realtime audio, configurable reasoning, precise spoken-information handling, and function calling. Its 128K context window, 32K maximum output, image input, and text-and-audio output support complex conversational systems. The main costs are audio usage, integration complexity, and the absence of video, structured outputs, fine-tuning, embeddings, and native image or video generation. It is a strong fit when natural, tool-enabled voice interaction is central to the application, but a specialized text or multimodal model may be more efficient for non-realtime workloads.

