GPT-Realtime

GPT-Realtime

by OpenAI · Deprecated; scheduled for API shutdown on January 20, 2027

OpenAI GPT-Realtime is a low-latency model for speech-to-speech and multimodal applications. It supports text and audio in both directions, image input, function calling, and WebRTC, WebSocket, and SIP connections. Its 32,000-token context window and separate audio pricing suit realtime voice agents, but the deprecated model family is scheduled for API shutdown on January 20, 2027.

Text Speech Reasoning Coding
GPT-Realtime is OpenAI’s first generally available realtime model for production voice agents. It is designed for conversations in which users speak naturally, interrupt the assistant, and expect an immediate response rather than waiting for a separate speech-recognition, text-generation, and speech-synthesis pipeline. The model accepts text, audio, and image inputs and can return text or audio. It also supports function calling for connecting a voice agent to application-specific tools. However, OpenAI has classified the model family as deprecated and scheduled it for API shutdown on January 20, 2027, so existing integrations and short-term projects are its strongest fit.
Outputs

What GPT-Realtime can produce

Text Speech
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Tool use Prompt caching Batch API Multimodal output
Model profile

Performance characteristics

5/10 Reasoning
4/10 Coding
9/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime
Model type Realtime
Context window 32K tokens
Maximum output 4K tokens
Knowledge cutoff 2023-10-01
Release date 2025-08-28
Status Deprecated; scheduled for API shutdown on January 20, 2027
Deprecation date 2026-07-20
Shutdown date 2027-01-20
Knowledge cutoff notes

The official model page lists October 1, 2023 as the knowledge cutoff. Realtime web search or external tools do not change the underlying model knowledge cutoff.

Model notes

Canonical model ID is gpt-realtime. OpenAI describes it as its first generally available realtime model. It supports text and audio inputs and outputs, image input, function calling, and realtime connections over WebRTC, WebSocket, or SIP. The model page lists a 32,000-token context window, 4,096 maximum output tokens, and an October 1, 2023 knowledge cutoff. OpenAI announced deprecation of the gpt-realtime model family on July 20, 2026, with API shutdown scheduled for January 20, 2027 and GPT-Realtime-2.1 listed as the recommended replacement. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input Text: $4.00 per 1M tokens; cached text: $0.40 per 1M tokens; audio: $32.00 per 1M tokens; cached audio: $0.40 per 1M tokens; image: $5.00 per 1M tokens; cached image: $0.50 per 1M tokens
Output Text: $16.00 per 1M tokens; audio: $64.00 per 1M tokens
Model guide

GPT-Realtime: OpenAI’s Low-Latency Speech-to-Speech Model

GPT-Realtime is OpenAI’s generally available realtime model for building low-latency voice and multimodal applications. It accepts text, audio, and images, produces text or natural-sounding audio, supports function calling, and connects through WebRTC, WebSocket, or SIP. Its 32,000-token context window and separate text and audio pricing make it suitable for realtime agents, although its planned API shutdown makes it a poor choice for new long-term deployments.

What is GPT-Realtime?

GPT-Realtime is a realtime multimodal model from OpenAI. Its primary purpose is speech-to-speech interaction: a user can talk to an application, the model can interpret the audio, and it can answer with generated audio during an ongoing session. Text can also be used as an input or output, and images can be supplied as input for visual discussion.

Unlike a traditional voice application that separately transcribes speech, sends the transcript to a text model, and then passes the answer to a speech synthesizer, GPT-Realtime is designed for direct realtime interaction. This architecture is intended to reduce conversational delay and make natural turn-taking, streaming audio, and interruptions easier to manage.

OpenAI made GPT-Realtime generally available for production voice agents. It can be used for voice assistants, customer-support systems, education tools, accessibility interfaces, and other applications where a user needs to communicate with an AI system by speaking. The model’s canonical model identifier is gpt-realtime.

Where GPT-Realtime fits in OpenAI’s catalog

GPT-Realtime belongs to OpenAI’s realtime model family rather than serving as a general-purpose text model. Its distinguishing concern is low-latency, persistent interaction involving audio. The model documentation identifies it as a generally available realtime model, but OpenAI has since classified the gpt-realtime family as legacy and deprecated.

OpenAI announced the deprecation on July 20, 2026, and lists January 20, 2027 as the planned API shutdown date. GPT-Realtime-2.1 is identified as the recommended replacement in the supplied lifecycle information. That positioning matters when selecting the model: GPT-Realtime may still be useful for maintaining an existing integration or evaluating an established application, but a new production system expected to operate beyond the shutdown date should assess the recommended newer realtime option instead.

Inputs, outputs, and supported modalities

GPT-Realtime supports text and audio in both directions. Audio input allows an application to receive a user’s speech, while audio output allows the model to respond with generated speech. Text input and output remain available for instructions, transcripts, application messages, or interfaces that combine voice with written content.

The model also accepts images as input. For example, a voice assistant could discuss an image supplied during a session. Image input does not mean that GPT-Realtime generates images: the supplied specifications identify text and audio as its output modalities, with no image or video output.

  • Text input: Supported.
  • Audio input: Supported.
  • Image input: Supported.
  • Video input: Not supported as a native model modality.
  • Text output: Supported.
  • Audio output: Supported.
  • Image and video output: Not supported.

Realtime sessions are designed to preserve conversational state and handle events such as audio turns, user interruptions, and tool calls. These behaviors are particularly relevant to voice agents: the application does not have to treat every spoken utterance as an isolated request.

Connections and function calling

Applications can connect to GPT-Realtime through WebRTC, WebSocket, or SIP. WebRTC is commonly associated with browser and interactive communications, WebSocket provides a persistent application connection, and SIP supports telephony-oriented integrations. The appropriate transport depends on the environment in which the voice agent operates.

GPT-Realtime supports function calling, which lets the model request an action from an external application. A customer-support agent, for example, could call a company system to retrieve an account detail, while an education application could invoke a tool that looks up course information. The external application remains responsible for defining the available functions, executing them, and returning the results.

Function calling should not be confused with built-in web search. The supplied model documentation does not list native web search support, and it also does not list structured-output support. Developers who need current external information or strict machine-readable response formats would need to evaluate another model or implement an appropriate application-side design.

Context window and technical limits

The verified context window is 32,000 tokens, and the maximum output is 4,096 tokens. A context window is the amount of conversation and other model-visible information that can be considered during a request or session; it is not the same as the duration of an audio call. The supplied research does not provide a separate maximum call duration or audio-duration limit, so those values should not be inferred from the token limits.

SpecificationGPT-Realtime
Canonical model IDgpt-realtime
ProviderOpenAI
Model typeRealtime
Context window32,000 tokens
Maximum output4,096 tokens
Knowledge cutoffOctober 1, 2023
Connection methodsWebRTC, WebSocket, and SIP
Function callingSupported
Native web searchNot supported according to the supplied model documentation
Structured outputsNot supported according to the supplied model documentation

The October 1, 2023 knowledge cutoff is the underlying model cutoff. Connecting the model to an external tool can provide newer information, but it does not update the model’s built-in knowledge.

GPT-Realtime pricing

GPT-Realtime uses separate prices for text and audio tokens. Audio is substantially more expensive than text, so the cost profile depends heavily on how much spoken input and output a session generates. Cached input prices are lower than the corresponding uncached input prices listed for text and audio.

Usage typePrice per 1 million tokens
Text input$4.00
Cached text input$0.40
Text output$16.00
Audio input$32.00
Cached audio input$0.40
Audio output$64.00
Image input$5.00
Cached image input$0.50

These are token prices rather than a fixed per-minute voice rate. Actual spending depends on the amount and type of input and output, how much context is retained, and how often the application uses audio versus text. In particular, applications that produce long spoken answers can incur more output cost than text-oriented applications because the listed audio output price is higher than the text output price.

Strengths and trade-offs

GPT-Realtime’s main strength is the combination of direct audio interaction, low-latency session behavior, and tool use. It is a focused option for applications where conversational timing matters more than maximum reasoning depth. A voice agent can listen, respond in audio, handle interruptions, and call external functions without requiring the developer to assemble a separate transcription and speech-generation pipeline around a standard text model.

Its multimodal input is another practical advantage. An application can combine spoken questions with an image, allowing use cases such as discussing visual material through a voice interface. The model also supports several connection methods, which gives developers options for browser, server, and telephony-oriented deployments.

The trade-offs are equally important. The model does not provide native web search or structured outputs according to the supplied documentation. Its knowledge cutoff is October 1, 2023, so current information must come from external tools. Audio output costs $64 per 1 million tokens, while text output costs $16 per 1 million tokens, making voice-heavy applications more expensive than equivalent text interactions. The model’s editorial scores in the supplied data rate speed highly at 9 out of 10, while reasoning is rated 5 out of 10, coding 4 out of 10, and cost efficiency 6 out of 10. These are comparative editorial estimates, not OpenAI benchmarks or provider-published guarantees.

When to choose GPT-Realtime

GPT-Realtime is most appropriate when realtime spoken interaction is the central requirement and the project is already built around the model or has a limited deployment horizon. Suitable examples include:

  • Voice assistants that need rapid spoken responses and interruption handling.
  • Customer-support agents that retrieve information or perform actions through function calls.
  • Education and accessibility applications where users interact primarily by voice.
  • Conversational interfaces that need both audio and image input.
  • Existing integrations that need continued operation before the planned shutdown date.

It is less suitable for a new long-lived production deployment because OpenAI has scheduled the model family for API shutdown on January 20, 2027. A successor or newer realtime model should be evaluated when long-term availability is a requirement. Another option may also be more appropriate when the application depends on advanced reasoning, strong coding performance, native web search, strict structured outputs, video understanding, or lower audio costs.

Practical evaluation before deployment

Teams considering GPT-Realtime should test the complete interaction rather than judging the model only from text responses. Measure perceived response delay, interruption behavior, audio quality, tool-call reliability, and the cost of typical spoken sessions. Include realistic context lengths because the 32,000-token window must accommodate instructions, conversation history, tool results, and any supplied visual or textual material.

It is also important to design for lifecycle risk. Existing users should confirm how their integration will migrate before January 20, 2027. New users should compare the recommended successor against GPT-Realtime on the exact requirements that matter: supported transports, audio behavior, function calling, pricing, context limits, and availability after the announced shutdown date.

Bottom line

GPT-Realtime is a specialized OpenAI model for fast speech-to-speech and multimodal sessions. It combines text and audio input and output, image input, persistent realtime connections, and function calling in a single model-oriented workflow. Those capabilities make it a practical fit for voice agents and conversational applications, while its pricing, limited web and structured-output support, modest editorial reasoning and coding ratings, and scheduled API shutdown limit its appeal for new long-term systems.


Answers to Frequently Asked Questions

Is GPT-Realtime still suitable for new production applications?
GPT-Realtime may be suitable for existing integrations or projects with a limited deployment horizon, but OpenAI has classified the gpt-realtime family as legacy and deprecated. The planned API shutdown date is January 20, 2027, and GPT-Realtime-2.1 is identified as the recommended replacement for longer-term deployments.
Does GPT-Realtime support function calling and web search?
GPT-Realtime supports function calling, allowing it to request actions from external applications such as retrieving account information. According to the supplied model documentation, it does not provide native web search or structured-output support, so developers must use application-side tools or consider another model for those requirements.
How can applications connect to GPT-Realtime?
Applications can connect to GPT-Realtime through WebRTC, WebSocket, or SIP. WebRTC is suited to browser and interactive communications, WebSocket provides a persistent application connection, and SIP supports telephony integrations.
What is GPT-Realtime used for?
GPT-Realtime is OpenAI’s low-latency multimodal model for realtime speech-to-speech interaction. It is designed for voice assistants, customer-support agents, education tools, accessibility interfaces, and other applications that require spoken conversations, interruptions, persistent sessions, and tool use.
What inputs and outputs does GPT-Realtime support?
GPT-Realtime supports text and audio input and output, as well as image input. It does not support native video input or image and video output. Applications can combine spoken questions with images for visual discussion.


Sources 5
Provider

About OpenAI