GPT-Realtime

GPT-Realtime-2

by OpenAI · Current

OpenAI’s GPT-Realtime-2 is a reasoning-focused voice model for low-latency speech-to-speech applications. It supports text, image, and audio input, text and audio output, configurable reasoning effort, tool calling, and a 128,000-token context window.

Text Speech Reasoning Coding
GPT-Realtime-2 is designed for production voice agents that need more than fast conversational responses. The model combines speech-to-speech interaction with GPT-5-class reasoning, stronger instruction following, configurable reasoning effort, long-session context, and tool use for workflows such as customer support, scheduling, transaction assistance, and live voice automation.
Outputs

What GPT-Realtime-2 can produce

Text Speech
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Tool use Streaming Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

9/10 Reasoning
7/10 Coding
7/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime
Model type Multimodal
Context window 128K tokens
Maximum output 32K tokens
Knowledge cutoff 2024-09-30
Release date 2026-05-07
Status Current
Knowledge cutoff notes

The official model documentation lists September 30, 2024 as the model's knowledge cutoff. Web search or external tools used during an application session do not change the underlying cutoff.

Model notes

GPT-Realtime-2 is OpenAI's reasoning-focused realtime voice model and was announced on May 7, 2026. It supports configurable reasoning effort, stronger instruction following, tool calling, text and audio output, and text, image, and audio input. The official model page lists general streaming as not supported while the model's core Realtime API usage is designed for low-latency streamed speech-to-speech sessions; endpoint-specific behavior should be verified. Structured outputs are listed as unsupported, and a separate legacy JSON-mode capability was not verified. Higher reasoning effort can increase latency and output-token usage. The model's knowledge cutoff is September 30, 2024.

Cost

Model pricing

Input Text: $4.00 per 1M tokens; cached text: $0.40 per 1M; audio: $32.00 per 1M tokens; cached audio: $0.40 per 1M; image: $5.00 per 1M tokens; cached image: $0.50 per 1M
Output Text: $24.00 per 1M tokens; audio: $64.00 per 1M tokens
Model guide

GPT-Realtime-2: Features, Pricing, Context and API Support

GPT-Realtime-2 is OpenAI’s reasoning-focused realtime voice model for low-latency speech-to-speech applications. It accepts text, images, and audio; produces text and audio; supports configurable reasoning effort and tool calling; and provides a 128,000-token context window for longer voice-agent sessions.

What is GPT-Realtime-2?

GPT-Realtime-2 is OpenAI’s reasoning-focused realtime voice model for low-latency speech-to-speech applications. Instead of requiring developers to connect separate speech recognition, language reasoning, and text-to-speech systems, it is designed to listen, reason, speak, and call tools within one live interaction.

OpenAI introduced GPT-Realtime-2 on May 7, 2026, describing it as its most capable realtime voice model and its first voice model with GPT-5-class reasoning. It sits in OpenAI’s realtime model lineup, aimed at applications where the system must respond naturally while also following instructions, considering context, and taking approved actions.

The model is available through OpenAI’s Realtime API and related API surfaces. Its primary audience is developers building persistent voice experiences rather than users looking for a general-purpose chat model or an image and video generation system.

Supported inputs and outputs

GPT-Realtime-2 accepts text, images, and audio. It can return text and audio, allowing an application to provide spoken responses while also retaining transcripts, text instructions, or machine-readable conversational state where supported by the surrounding application.

CapabilitySupport
Text inputSupported
Audio inputSupported
Image inputSupported
Video inputNot supported
Text outputSupported
Audio outputSupported
Image or video outputNot supported
Tool and function callingSupported

Image input can provide visual context during a voice session, such as a document, product photo, or screenshot, but GPT-Realtime-2 is not an image-generation model. Likewise, it can conduct spoken conversations but does not generate video.

Reasoning, context and output limits

GPT-Realtime-2 has a 128,000-token context window and a maximum output of 32,000 tokens. The context window is the amount of conversation, instructions, retrieved information, and other material the model can consider in a session. This capacity is useful for long-running voice agents that need to retain policies, account notes, conversation history, or operational records.

The model supports configurable reasoning effort. In practical terms, developers can choose how much deliberation the model should apply before responding or calling a tool. Higher reasoning effort may improve decision quality for complicated workflows, but it can also increase latency and output-token usage. That trade-off matters more in voice applications than in ordinary text chat because users notice pauses immediately.

GPT-Realtime-2 also supports brief spoken preambles before longer reasoning or tool-use operations. A voice agent can therefore acknowledge a request before completing a slower lookup or multi-step decision, helping the interaction feel less interrupted while the model works.

How GPT-Realtime-2 handles voice-agent workflows

The model is intended for agents that must do more than answer questions. Through tool calling, it can be connected to application functions such as checking an order, retrieving account information, scheduling an appointment, or initiating an approved business workflow.

Tool calling does not mean the model should receive unrestricted control over business systems. Developers should define which tools are available, what arguments are valid, which actions require confirmation, and when the agent must transfer the conversation to a person. For example, an agent might retrieve an account balance automatically but require explicit confirmation before changing billing details or submitting a transaction.

GPT-Realtime-2’s reasoning capability is especially relevant when a request requires several steps. The model can interpret the user’s request, identify missing information, choose an appropriate function, and explain the result through speech. However, the application remains responsible for authentication, authorization, validation, audit logging, and safeguards around consequential actions.

GPT-Realtime-2 pricing

OpenAI prices GPT-Realtime-2 by token type rather than using one simple per-request rate. Text, audio, and image tokens have separate prices, and cached input is cheaper than uncached input. The listed rates are:

Token categoryPrice per 1 million tokens
Text input$4.00
Cached text input$0.40
Text output$24.00
Audio input$32.00
Cached audio input$0.40
Audio output$64.00
Image input$5.00
Cached image input$0.50

Audio output is substantially more expensive than text output, so applications should avoid generating unnecessary spoken content. Long system prompts, repeated instructions, and recurring context may also affect cost, although cached input pricing can reduce the cost of reused material when caching applies.

The actual cost of a voice session depends on its balance of input and output audio, text instructions, images, cached content, reasoning effort, and conversation length. Teams evaluating the model should measure representative sessions rather than estimating expenses from text-token pricing alone.

API access and realtime behavior

GPT-Realtime-2 is available through OpenAI’s Realtime API, which is designed for persistent, low-latency voice sessions over supported realtime transports. This is the API path to consider when the application needs to stream audio during an ongoing conversation.

The supplied model catalog separately lists general streaming support as unavailable for some non-realtime API behavior, while the model’s core Realtime API usage is built around streamed speech-to-speech interaction. Developers should therefore verify the exact endpoint, transport, and response mode used by their implementation instead of assuming that every API surface has identical streaming behavior.

The model catalog lists tool use, caching, and batch API support. Structured outputs are listed as unsupported, and a separate legacy JSON-mode capability has not been verified. Applications that require strict machine-readable JSON should not assume that GPT-Realtime-2 can guarantee a schema-compliant response simply because it can return text.

Main strengths and limitations

GPT-Realtime-2’s main strength is the combination of realtime voice interaction with deeper reasoning. A basic speech-to-speech system may be well suited to short conversational exchanges, but GPT-Realtime-2 is aimed at sessions where the agent must interpret policies, use background information, ask for missing details, and perform tool-assisted tasks.

  • Reasoning in live conversations: Configurable reasoning effort supports more involved decisions than a simple response-generation workflow.
  • Long-session context: The 128,000-token context window accommodates substantial instructions, records, and conversation history.
  • Multimodal context: Text, audio, and image input allow a voice agent to work with more than spoken words.
  • Natural audio interaction: Audio output supports direct speech-to-speech experiences without a separate text-to-speech stage in the main workflow.
  • Action-oriented workflows: Tool calling connects the model to approved business functions and external data.

There are also important limitations. GPT-Realtime-2 does not generate images or video, does not support video input, and is not suited to applications that require structured outputs as a verified native capability. Its audio-token prices are higher than its text-token prices, and higher reasoning settings can increase response latency.

It is also not automatically the best option for every voice application. If the task consists mainly of short, predictable exchanges, a less expensive and faster non-reasoning realtime model may be more economical. If the main requirement is strict JSON generation, image generation, video generation, embeddings, or fine-tuning, another model type is more appropriate.

Best use cases

GPT-Realtime-2 is a good fit when spoken interaction and reliable multi-step behavior must occur together. Suitable applications include:

  • Customer-support and contact-center agents that retrieve records or follow policy workflows.
  • Scheduling assistants that gather details, check availability, and create appointments through tools.
  • Voice-based account assistants that explain information and request confirmation before sensitive actions.
  • Accessibility interfaces that let users interact with software through natural speech.
  • Live assistants that combine spoken conversation with image or document context.
  • Education and coaching experiences that need longer conversational continuity.

For each use case, the application should distinguish between information retrieval and actions that change data. The model can help decide what to do, but the surrounding software should enforce permissions and confirmation requirements.

When to choose GPT-Realtime-2

Choose GPT-Realtime-2 when the defining requirement is a realtime voice agent that can reason, remember substantial session context, interpret text or images, and use tools. It is particularly compelling when a short answer is not enough and the agent must complete a workflow during the conversation.

Choose a simpler realtime model when minimum latency and low cost matter more than complex reasoning. A conventional text model may be preferable when audio interaction is unnecessary. A specialized image or video model is the better choice for visual generation, while a model with verified structured-output support should be used when downstream software depends on strict schemas.

The central trade-off is capability versus responsiveness and cost. GPT-Realtime-2 offers more reasoning and workflow support than a basic voice bot, but its reasoning settings, audio pricing, and long-context processing should be justified by the application’s requirements. Teams should test real conversations, interruptions, tool failures, confirmation flows, and average audio duration before selecting it for production.

Bottom line

GPT-Realtime-2 is an OpenAI model for low-latency voice agents that need reasoning and tool use, not merely speech conversion. It accepts text, images, and audio, produces text and audio, supports a 128,000-token context window, and can reason before responding or taking an approved tool action. Its strongest use cases involve long, multimodal, action-oriented conversations. Its higher audio costs, possible reasoning latency, lack of structured-output support, and absence of image or video generation make a simpler or more specialized model a better choice for narrower requirements.


Answers to Frequently Asked Questions

When should developers choose GPT-Realtime-2?
Developers should choose GPT-Realtime-2 when they need a realtime voice agent that can reason through multi-step tasks, maintain substantial context, interpret images, and use tools such as scheduling or account-management functions. A simpler realtime model may be better when low cost and minimum latency are the top priorities.
How much does GPT-Realtime-2 cost?
GPT-Realtime-2 uses token-based pricing. Text input costs $4 per million tokens, text output $24, audio input $32, audio output $64, and image input $5. Cached input is less expensive, with prices ranging from $0.40 to $0.50 per million tokens depending on the modality.
What is GPT-Realtime-2’s context window and maximum output?
GPT-Realtime-2 has a 128,000-token context window and a maximum output of 32,000 tokens. It also supports configurable reasoning effort, which can improve complex decision-making but may increase latency and token usage.
What is GPT-Realtime-2 designed for?
GPT-Realtime-2 is OpenAI’s reasoning-focused realtime voice model for low-latency speech-to-speech applications. It can listen, reason, speak, and call approved tools within a live interaction.
What inputs and outputs does GPT-Realtime-2 support?
GPT-Realtime-2 accepts text, audio, and images, and produces text and audio. It does not support video input or image and video generation.


Sources 5
Provider

About OpenAI