What is GPT-Realtime-2?
GPT-Realtime-2 is OpenAI’s reasoning-focused realtime voice model for low-latency speech-to-speech applications. Instead of requiring developers to connect separate speech recognition, language reasoning, and text-to-speech systems, it is designed to listen, reason, speak, and call tools within one live interaction.
OpenAI introduced GPT-Realtime-2 on May 7, 2026, describing it as its most capable realtime voice model and its first voice model with GPT-5-class reasoning. It sits in OpenAI’s realtime model lineup, aimed at applications where the system must respond naturally while also following instructions, considering context, and taking approved actions.
The model is available through OpenAI’s Realtime API and related API surfaces. Its primary audience is developers building persistent voice experiences rather than users looking for a general-purpose chat model or an image and video generation system.
Supported inputs and outputs
GPT-Realtime-2 accepts text, images, and audio. It can return text and audio, allowing an application to provide spoken responses while also retaining transcripts, text instructions, or machine-readable conversational state where supported by the surrounding application.
| Capability | Support |
|---|---|
| Text input | Supported |
| Audio input | Supported |
| Image input | Supported |
| Video input | Not supported |
| Text output | Supported |
| Audio output | Supported |
| Image or video output | Not supported |
| Tool and function calling | Supported |
Image input can provide visual context during a voice session, such as a document, product photo, or screenshot, but GPT-Realtime-2 is not an image-generation model. Likewise, it can conduct spoken conversations but does not generate video.
Reasoning, context and output limits
GPT-Realtime-2 has a 128,000-token context window and a maximum output of 32,000 tokens. The context window is the amount of conversation, instructions, retrieved information, and other material the model can consider in a session. This capacity is useful for long-running voice agents that need to retain policies, account notes, conversation history, or operational records.
The model supports configurable reasoning effort. In practical terms, developers can choose how much deliberation the model should apply before responding or calling a tool. Higher reasoning effort may improve decision quality for complicated workflows, but it can also increase latency and output-token usage. That trade-off matters more in voice applications than in ordinary text chat because users notice pauses immediately.
GPT-Realtime-2 also supports brief spoken preambles before longer reasoning or tool-use operations. A voice agent can therefore acknowledge a request before completing a slower lookup or multi-step decision, helping the interaction feel less interrupted while the model works.
How GPT-Realtime-2 handles voice-agent workflows
The model is intended for agents that must do more than answer questions. Through tool calling, it can be connected to application functions such as checking an order, retrieving account information, scheduling an appointment, or initiating an approved business workflow.
Tool calling does not mean the model should receive unrestricted control over business systems. Developers should define which tools are available, what arguments are valid, which actions require confirmation, and when the agent must transfer the conversation to a person. For example, an agent might retrieve an account balance automatically but require explicit confirmation before changing billing details or submitting a transaction.
GPT-Realtime-2’s reasoning capability is especially relevant when a request requires several steps. The model can interpret the user’s request, identify missing information, choose an appropriate function, and explain the result through speech. However, the application remains responsible for authentication, authorization, validation, audit logging, and safeguards around consequential actions.
GPT-Realtime-2 pricing
OpenAI prices GPT-Realtime-2 by token type rather than using one simple per-request rate. Text, audio, and image tokens have separate prices, and cached input is cheaper than uncached input. The listed rates are:
| Token category | Price per 1 million tokens |
|---|---|
| Text input | $4.00 |
| Cached text input | $0.40 |
| Text output | $24.00 |
| Audio input | $32.00 |
| Cached audio input | $0.40 |
| Audio output | $64.00 |
| Image input | $5.00 |
| Cached image input | $0.50 |
Audio output is substantially more expensive than text output, so applications should avoid generating unnecessary spoken content. Long system prompts, repeated instructions, and recurring context may also affect cost, although cached input pricing can reduce the cost of reused material when caching applies.
The actual cost of a voice session depends on its balance of input and output audio, text instructions, images, cached content, reasoning effort, and conversation length. Teams evaluating the model should measure representative sessions rather than estimating expenses from text-token pricing alone.
API access and realtime behavior
GPT-Realtime-2 is available through OpenAI’s Realtime API, which is designed for persistent, low-latency voice sessions over supported realtime transports. This is the API path to consider when the application needs to stream audio during an ongoing conversation.
The supplied model catalog separately lists general streaming support as unavailable for some non-realtime API behavior, while the model’s core Realtime API usage is built around streamed speech-to-speech interaction. Developers should therefore verify the exact endpoint, transport, and response mode used by their implementation instead of assuming that every API surface has identical streaming behavior.
The model catalog lists tool use, caching, and batch API support. Structured outputs are listed as unsupported, and a separate legacy JSON-mode capability has not been verified. Applications that require strict machine-readable JSON should not assume that GPT-Realtime-2 can guarantee a schema-compliant response simply because it can return text.
Main strengths and limitations
GPT-Realtime-2’s main strength is the combination of realtime voice interaction with deeper reasoning. A basic speech-to-speech system may be well suited to short conversational exchanges, but GPT-Realtime-2 is aimed at sessions where the agent must interpret policies, use background information, ask for missing details, and perform tool-assisted tasks.
- Reasoning in live conversations: Configurable reasoning effort supports more involved decisions than a simple response-generation workflow.
- Long-session context: The 128,000-token context window accommodates substantial instructions, records, and conversation history.
- Multimodal context: Text, audio, and image input allow a voice agent to work with more than spoken words.
- Natural audio interaction: Audio output supports direct speech-to-speech experiences without a separate text-to-speech stage in the main workflow.
- Action-oriented workflows: Tool calling connects the model to approved business functions and external data.
There are also important limitations. GPT-Realtime-2 does not generate images or video, does not support video input, and is not suited to applications that require structured outputs as a verified native capability. Its audio-token prices are higher than its text-token prices, and higher reasoning settings can increase response latency.
It is also not automatically the best option for every voice application. If the task consists mainly of short, predictable exchanges, a less expensive and faster non-reasoning realtime model may be more economical. If the main requirement is strict JSON generation, image generation, video generation, embeddings, or fine-tuning, another model type is more appropriate.
Best use cases
GPT-Realtime-2 is a good fit when spoken interaction and reliable multi-step behavior must occur together. Suitable applications include:
- Customer-support and contact-center agents that retrieve records or follow policy workflows.
- Scheduling assistants that gather details, check availability, and create appointments through tools.
- Voice-based account assistants that explain information and request confirmation before sensitive actions.
- Accessibility interfaces that let users interact with software through natural speech.
- Live assistants that combine spoken conversation with image or document context.
- Education and coaching experiences that need longer conversational continuity.
For each use case, the application should distinguish between information retrieval and actions that change data. The model can help decide what to do, but the surrounding software should enforce permissions and confirmation requirements.
When to choose GPT-Realtime-2
Choose GPT-Realtime-2 when the defining requirement is a realtime voice agent that can reason, remember substantial session context, interpret text or images, and use tools. It is particularly compelling when a short answer is not enough and the agent must complete a workflow during the conversation.
Choose a simpler realtime model when minimum latency and low cost matter more than complex reasoning. A conventional text model may be preferable when audio interaction is unnecessary. A specialized image or video model is the better choice for visual generation, while a model with verified structured-output support should be used when downstream software depends on strict schemas.
The central trade-off is capability versus responsiveness and cost. GPT-Realtime-2 offers more reasoning and workflow support than a basic voice bot, but its reasoning settings, audio pricing, and long-context processing should be justified by the application’s requirements. Teams should test real conversations, interruptions, tool failures, confirmation flows, and average audio duration before selecting it for production.
Bottom line
GPT-Realtime-2 is an OpenAI model for low-latency voice agents that need reasoning and tool use, not merely speech conversion. It accepts text, images, and audio, produces text and audio, supports a 128,000-token context window, and can reason before responding or taking an approved tool action. Its strongest use cases involve long, multimodal, action-oriented conversations. Its higher audio costs, possible reasoning latency, lack of structured-output support, and absence of image or video generation make a simpler or more specialized model a better choice for narrower requirements.

