GPT-Realtime

GPT-Realtime-2.1

by OpenAI · current

OpenAI's GPT-Realtime-2.1 is a realtime reasoning model for voice agents. It supports text, audio and image input, text and audio output, function calling, configurable reasoning effort, a 128K context window and 32K maximum output. It is designed for low-latency speech-to-speech applications that need precise instruction following, improved interruption handling and external tool integration.

Text Speech Reasoning Coding
GPT-Realtime-2.1 is OpenAI's July 2026 update to GPT-Realtime-2, built for realtime voice agents rather than ordinary text-only chat. It combines speech-to-speech interaction, image understanding, configurable reasoning, and tool use in a single model. The main improvements are better recognition of alphanumeric information, improved handling of silence and background noise, and more natural interruption behavior. Its 128,000-token context window and 32,000-token maximum output make it suitable for complex conversational workflows, although audio usage is considerably more expensive than text-only processing.
Outputs

What GPT-Realtime-2.1 can produce

Text Speech
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Tool use Streaming Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

8/10 Reasoning
7/10 Coding
8/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime
Model type Realtime
Context window 128K tokens
Maximum output 32K tokens
Knowledge cutoff 2024-09-30
Release date 2026-07-06
Status current
Knowledge cutoff notes

OpenAI's model page lists September 30, 2024 as the knowledge cutoff. External tools or web-search integrations may provide newer information during use but do not change the underlying cutoff.

Model notes

GPT-Realtime-2.1 is OpenAI's July 6, 2026 update to GPT-Realtime-2. It improves alphanumeric recognition, silence and noise handling, and interruption behavior. It supports configurable reasoning effort; higher effort can increase latency and output-token usage. The model supports text, audio, and image input and text and audio output. It does not support video input, structured outputs, fine-tuning, image generation, video generation, or embeddings. The model page lists standard streaming as unsupported, while Realtime API sessions exchange streamed events and audio over WebRTC or WebSocket. The documented knowledge cutoff is September 30, 2024.

Cost

Model pricing

Input Text: $4.00 per 1M tokens; cached text: $0.40 per 1M; audio: $32.00 per 1M audio tokens; cached audio: $0.40 per 1M; image: $5.00 per 1M tokens; cached image: $0.50 per 1M
Output Text: $24.00 per 1M tokens; audio: $64.00 per 1M audio tokens
Model guide

GPT-Realtime-2.1: Features, Pricing, Limits and API Support

GPT-Realtime-2.1 is OpenAI's realtime reasoning model for low-latency speech-to-speech applications. It accepts text, audio, and image input; produces text and audio; supports configurable reasoning effort and function calling; and is designed for voice agents that need accurate spoken information handling, external tools, and natural interruption behavior.

What is GPT-Realtime-2.1?

GPT-Realtime-2.1 is OpenAI's current realtime reasoning model for interactive voice and multimodal applications. Unlike a conventional text model that receives a prompt and returns a completion, it is designed to maintain a live conversation in which audio can be exchanged continuously and the user can interrupt the assistant.

The model can understand text, speech, and images. It can respond with text or synthesized audio, allowing developers to build speech-to-speech experiences without treating transcription and speech generation as completely separate application steps. Function calling lets it invoke external business logic, such as looking up an account, checking an order, scheduling an appointment, or updating a customer record.

OpenAI released GPT-Realtime-2.1 on July 6, 2026, as an update to GPT-Realtime-2. Its primary role in OpenAI's model lineup is realtime reasoning for voice agents, rather than inexpensive bulk text generation, image creation, video understanding, or general-purpose embeddings.

What changed from GPT-Realtime-2?

According to the supplied OpenAI documentation, GPT-Realtime-2.1 improves several parts of voice interaction:

  • Alphanumeric recognition: better handling of order numbers, account identifiers, confirmation codes, phone numbers, addresses, and similar information spoken aloud.
  • Silence and background-noise handling: improved behavior when a caller pauses, speaks in a noisy environment, or takes time to respond.
  • Interruption behavior: more natural handling of users who begin speaking while the assistant is still responding.
  • Configurable reasoning effort: developers can adjust how much reasoning the model applies. Higher effort may help with difficult decisions and tool-use flows, but can increase latency and output-token consumption.

These improvements matter most in production voice applications, where a misheard identifier or poorly handled interruption can be more damaging than a slightly imperfect response in a casual conversation.

Capabilities and supported modalities

GPT-Realtime-2.1 supports multiple input and output types, but it is not a model for every kind of multimodal generation. Its supported modalities are:

CapabilitySupport
Text inputYes
Audio inputYes
Image inputYes
Text outputYes
Audio outputYes
Video inputNo
Image generationNo
Video generationNo

Audio input and output enable direct speech-to-speech conversations. Image input can be useful when a voice agent needs to discuss a document, photograph, diagram, or other visual supplied by the user. The model does not natively create images or video, and it does not accept video input.

Context window and output limits

GPT-Realtime-2.1 has a documented context window of 128,000 tokens and a maximum output of 32,000 tokens. The context window is the amount of conversation and other input the model can consider in a request or realtime session. In a long-running voice application, developers still need to manage conversation history because retaining every turn increases input processing and cost.

The maximum output is substantially larger than most spoken responses require. In practice, voice applications generally need short, timely turns rather than tens of thousands of generated tokens. The limit is more relevant to complex tool workflows, detailed text responses, or applications that use the model for both conversational and written output.

Reasoning, coding, and tool use

GPT-Realtime-2.1 supports configurable reasoning effort. This gives developers a choice between faster responses for routine dialogue and more deliberate processing for difficult decisions, multi-step instructions, or tool calls. The trade-off is not simply quality versus speed: higher reasoning effort can also increase output-token usage and therefore operating cost.

Function calling is supported. In a voice agent, a function is an application-defined operation that the model can request, while the application performs the actual task and returns the result. Examples include retrieving a customer record, checking inventory, calculating a quote, or creating a calendar appointment. The model can then explain the result in the ongoing conversation.

The supplied research rates its coding capability as a moderate-to-strong editorial assessment rather than a provider-published benchmark. GPT-Realtime-2.1 can participate in coding-related tool workflows and generate text, but it is primarily optimized for realtime voice interaction. Teams focused on large codebases, long coding sessions, or text-only software development may prefer a model designed specifically for those workloads.

How the Realtime API fits in

GPT-Realtime-2.1 is used with OpenAI's Realtime API for stateful conversational sessions. The model identifier is gpt-realtime-2.1. Realtime sessions can use WebRTC or WebSocket transports, with the application managing session configuration, audio exchange, events, interruptions, and tool results.

WebRTC is commonly associated with interactive browser or client audio connections, while WebSocket connections are useful for server-managed realtime communication. The exact transport choice depends on the application architecture. In either case, the model's realtime behavior is based on streamed events and audio exchange rather than a single request followed by one final response.

Developers should distinguish realtime event streaming from the model's listed feature support. The supplied model information says standard streaming is not listed as a separate supported feature, while Realtime API sessions themselves exchange streamed events and audio. GPT-Realtime-2.1 also does not support structured outputs, so applications that require guaranteed schema-conforming JSON should consider another model or add their own validation and recovery layer.

GPT-Realtime-2.1 pricing

GPT-Realtime-2.1 uses usage-based pricing that varies by modality. The supplied OpenAI pricing information lists the following rates per 1 million units:

Usage typePrice
Text input$4.00 per 1 million tokens
Cached text input$0.40 per 1 million tokens
Text output$24.00 per 1 million tokens
Audio input$32.00 per 1 million audio tokens
Cached audio input$0.40 per 1 million audio tokens
Audio output$64.00 per 1 million audio tokens
Image input$5.00 per 1 million tokens
Cached image input$0.50 per 1 million tokens

Audio output is the most expensive listed output modality, and audio input is also much more expensive than ordinary text input. This makes the model a better fit for applications where realtime speech is central to the product than for high-volume text processing. Caching repeated context and limiting unnecessary conversation history can reduce recurring input costs.

Main strengths and trade-offs

GPT-Realtime-2.1's main strength is the combination of realtime audio, reasoning, multimodal input, and external tool use. A customer-service agent can listen to a caller, identify a spoken account number, consult a business system, and respond naturally without forcing the user through a separate transcription or menu-driven workflow.

Its improvements to silence, noise, alphanumeric recognition, and interruption behavior are particularly useful in telephony and contact-center environments. The model is also more suitable than a basic speech interface when the conversation requires decisions, policy interpretation, or several tool calls.

The principal trade-off is cost and complexity. Audio processing costs more than text processing, and realtime applications must handle transport connections, session state, interruptions, audio events, and tool results. Higher reasoning effort can improve difficult workflows but may make responses slower and more expensive. These costs should be evaluated against the value of immediate spoken interaction.

Best use cases

  • Customer-service and contact-center agents: handle spoken requests, retrieve records, and provide responses while maintaining a natural conversation.
  • Telephony and SIP-based systems: support callers who may speak over the assistant, pause frequently, or provide numbers and codes in noisy conditions.
  • Voice-operated business applications: connect speech interaction to scheduling, order management, account systems, or internal workflows.
  • Realtime support assistants: combine conversation with external information retrieval and action-taking tools.
  • Multimodal voice assistants: let users discuss an image, document, or other visual while speaking with the assistant.
  • Complex conversational workflows: use configurable reasoning when the assistant must interpret instructions before selecting a tool or response.

Limitations to consider

GPT-Realtime-2.1 does not support video input, image generation, video generation, structured outputs, fine-tuning, native embeddings, or web search as a listed built-in capability. It is therefore not a complete choice for applications that need visual video analysis, guaranteed JSON responses, custom model training, or vector embedding generation.

The documented knowledge cutoff is September 30, 2024. Connecting the model to external tools can provide newer information, but tool access does not change the model's underlying training cutoff. Applications that answer questions about current events, prices, inventory, policies, or account data should use an appropriate live data source and should not rely on model memory alone.

The model's audio pricing also makes it less attractive for inexpensive, high-volume, text-only workloads. A text-focused model may be a better choice when low cost and throughput matter more than natural voice interaction.

When to choose GPT-Realtime-2.1

Choose GPT-Realtime-2.1 when the product needs a live voice conversation and the assistant must do more than transcribe speech or read back scripted responses. It is especially appropriate when the system must recognize precise spoken information, tolerate interruptions, reason about a request, and call external tools during the same interaction.

Consider another type of model when the task is primarily text generation, inexpensive classification, embeddings, structured data production, image generation, or video understanding. A specialized text or coding model may offer a better speed-and-cost profile for those workloads. Likewise, if the application only needs speech recognition and speech synthesis with minimal reasoning, a simpler pipeline may be easier to operate and less expensive.

For teams comparing OpenAI model options, [GPT-4o](/openai/models/gpt-4o) can provide useful context for a general multimodal model, while [GPT-5.4](/openai/models/gpt-5-4) is a more relevant comparison for applications prioritizing general reasoning and text-based capability rather than realtime speech. These alternatives should not be treated as drop-in replacements: GPT-Realtime-2.1's distinguishing value is its realtime voice-agent workflow.

Bottom line

GPT-Realtime-2.1 is designed for production voice agents that need realtime audio, configurable reasoning, precise spoken-information handling, and function calling. Its 128K context window, 32K maximum output, image input, and text-and-audio output support complex conversational systems. The main costs are audio usage, integration complexity, and the absence of video, structured outputs, fine-tuning, embeddings, and native image or video generation. It is a strong fit when natural, tool-enabled voice interaction is central to the application, but a specialized text or multimodal model may be more efficient for non-realtime workloads.


Answers to Frequently Asked Questions

Does GPT-Realtime-2.1 support function calling and the Realtime API?
Yes. GPT-Realtime-2.1 uses the Realtime API with the model identifier gpt-realtime-2.1 and supports WebRTC and WebSocket connections. Function calling allows it to request actions such as checking orders, retrieving customer records, scheduling appointments, or updating business systems.
What are the context window and output limits of GPT-Realtime-2.1?
GPT-Realtime-2.1 has a 128,000-token context window and a maximum output limit of 32,000 tokens. Developers should still manage conversation history in long-running realtime sessions to control processing and costs.
How much does GPT-Realtime-2.1 cost?
GPT-Realtime-2.1 pricing varies by modality: text input costs $4 per 1 million tokens, text output $24, audio input $32 per 1 million audio tokens, audio output $64, and image input $5. Cached input rates are lower.
What is GPT-Realtime-2.1 designed for?
GPT-Realtime-2.1 is designed for realtime voice and multimodal applications, including voice agents, contact-center systems, telephony, and conversational assistants that need audio interaction, reasoning, and function calling.
What modalities does GPT-Realtime-2.1 support?
GPT-Realtime-2.1 supports text, audio, and image input, as well as text and audio output. It does not support video input, image generation, or video generation.


Sources 6
Provider

About OpenAI