What is GPT-Realtime?
GPT-Realtime is a realtime multimodal model from OpenAI. Its primary purpose is speech-to-speech interaction: a user can talk to an application, the model can interpret the audio, and it can answer with generated audio during an ongoing session. Text can also be used as an input or output, and images can be supplied as input for visual discussion.
Unlike a traditional voice application that separately transcribes speech, sends the transcript to a text model, and then passes the answer to a speech synthesizer, GPT-Realtime is designed for direct realtime interaction. This architecture is intended to reduce conversational delay and make natural turn-taking, streaming audio, and interruptions easier to manage.
OpenAI made GPT-Realtime generally available for production voice agents. It can be used for voice assistants, customer-support systems, education tools, accessibility interfaces, and other applications where a user needs to communicate with an AI system by speaking. The model’s canonical model identifier is gpt-realtime.
Where GPT-Realtime fits in OpenAI’s catalog
GPT-Realtime belongs to OpenAI’s realtime model family rather than serving as a general-purpose text model. Its distinguishing concern is low-latency, persistent interaction involving audio. The model documentation identifies it as a generally available realtime model, but OpenAI has since classified the gpt-realtime family as legacy and deprecated.
OpenAI announced the deprecation on July 20, 2026, and lists January 20, 2027 as the planned API shutdown date. GPT-Realtime-2.1 is identified as the recommended replacement in the supplied lifecycle information. That positioning matters when selecting the model: GPT-Realtime may still be useful for maintaining an existing integration or evaluating an established application, but a new production system expected to operate beyond the shutdown date should assess the recommended newer realtime option instead.
Inputs, outputs, and supported modalities
GPT-Realtime supports text and audio in both directions. Audio input allows an application to receive a user’s speech, while audio output allows the model to respond with generated speech. Text input and output remain available for instructions, transcripts, application messages, or interfaces that combine voice with written content.
The model also accepts images as input. For example, a voice assistant could discuss an image supplied during a session. Image input does not mean that GPT-Realtime generates images: the supplied specifications identify text and audio as its output modalities, with no image or video output.
- Text input: Supported.
- Audio input: Supported.
- Image input: Supported.
- Video input: Not supported as a native model modality.
- Text output: Supported.
- Audio output: Supported.
- Image and video output: Not supported.
Realtime sessions are designed to preserve conversational state and handle events such as audio turns, user interruptions, and tool calls. These behaviors are particularly relevant to voice agents: the application does not have to treat every spoken utterance as an isolated request.
Connections and function calling
Applications can connect to GPT-Realtime through WebRTC, WebSocket, or SIP. WebRTC is commonly associated with browser and interactive communications, WebSocket provides a persistent application connection, and SIP supports telephony-oriented integrations. The appropriate transport depends on the environment in which the voice agent operates.
GPT-Realtime supports function calling, which lets the model request an action from an external application. A customer-support agent, for example, could call a company system to retrieve an account detail, while an education application could invoke a tool that looks up course information. The external application remains responsible for defining the available functions, executing them, and returning the results.
Function calling should not be confused with built-in web search. The supplied model documentation does not list native web search support, and it also does not list structured-output support. Developers who need current external information or strict machine-readable response formats would need to evaluate another model or implement an appropriate application-side design.
Context window and technical limits
The verified context window is 32,000 tokens, and the maximum output is 4,096 tokens. A context window is the amount of conversation and other model-visible information that can be considered during a request or session; it is not the same as the duration of an audio call. The supplied research does not provide a separate maximum call duration or audio-duration limit, so those values should not be inferred from the token limits.
| Specification | GPT-Realtime |
|---|---|
| Canonical model ID | gpt-realtime |
| Provider | OpenAI |
| Model type | Realtime |
| Context window | 32,000 tokens |
| Maximum output | 4,096 tokens |
| Knowledge cutoff | October 1, 2023 |
| Connection methods | WebRTC, WebSocket, and SIP |
| Function calling | Supported |
| Native web search | Not supported according to the supplied model documentation |
| Structured outputs | Not supported according to the supplied model documentation |
The October 1, 2023 knowledge cutoff is the underlying model cutoff. Connecting the model to an external tool can provide newer information, but it does not update the model’s built-in knowledge.
GPT-Realtime pricing
GPT-Realtime uses separate prices for text and audio tokens. Audio is substantially more expensive than text, so the cost profile depends heavily on how much spoken input and output a session generates. Cached input prices are lower than the corresponding uncached input prices listed for text and audio.
| Usage type | Price per 1 million tokens |
|---|---|
| Text input | $4.00 |
| Cached text input | $0.40 |
| Text output | $16.00 |
| Audio input | $32.00 |
| Cached audio input | $0.40 |
| Audio output | $64.00 |
| Image input | $5.00 |
| Cached image input | $0.50 |
These are token prices rather than a fixed per-minute voice rate. Actual spending depends on the amount and type of input and output, how much context is retained, and how often the application uses audio versus text. In particular, applications that produce long spoken answers can incur more output cost than text-oriented applications because the listed audio output price is higher than the text output price.
Strengths and trade-offs
GPT-Realtime’s main strength is the combination of direct audio interaction, low-latency session behavior, and tool use. It is a focused option for applications where conversational timing matters more than maximum reasoning depth. A voice agent can listen, respond in audio, handle interruptions, and call external functions without requiring the developer to assemble a separate transcription and speech-generation pipeline around a standard text model.
Its multimodal input is another practical advantage. An application can combine spoken questions with an image, allowing use cases such as discussing visual material through a voice interface. The model also supports several connection methods, which gives developers options for browser, server, and telephony-oriented deployments.
The trade-offs are equally important. The model does not provide native web search or structured outputs according to the supplied documentation. Its knowledge cutoff is October 1, 2023, so current information must come from external tools. Audio output costs $64 per 1 million tokens, while text output costs $16 per 1 million tokens, making voice-heavy applications more expensive than equivalent text interactions. The model’s editorial scores in the supplied data rate speed highly at 9 out of 10, while reasoning is rated 5 out of 10, coding 4 out of 10, and cost efficiency 6 out of 10. These are comparative editorial estimates, not OpenAI benchmarks or provider-published guarantees.
When to choose GPT-Realtime
GPT-Realtime is most appropriate when realtime spoken interaction is the central requirement and the project is already built around the model or has a limited deployment horizon. Suitable examples include:
- Voice assistants that need rapid spoken responses and interruption handling.
- Customer-support agents that retrieve information or perform actions through function calls.
- Education and accessibility applications where users interact primarily by voice.
- Conversational interfaces that need both audio and image input.
- Existing integrations that need continued operation before the planned shutdown date.
It is less suitable for a new long-lived production deployment because OpenAI has scheduled the model family for API shutdown on January 20, 2027. A successor or newer realtime model should be evaluated when long-term availability is a requirement. Another option may also be more appropriate when the application depends on advanced reasoning, strong coding performance, native web search, strict structured outputs, video understanding, or lower audio costs.
Practical evaluation before deployment
Teams considering GPT-Realtime should test the complete interaction rather than judging the model only from text responses. Measure perceived response delay, interruption behavior, audio quality, tool-call reliability, and the cost of typical spoken sessions. Include realistic context lengths because the 32,000-token window must accommodate instructions, conversation history, tool results, and any supplied visual or textual material.
It is also important to design for lifecycle risk. Existing users should confirm how their integration will migrate before January 20, 2027. New users should compare the recommended successor against GPT-Realtime on the exact requirements that matter: supported transports, audio behavior, function calling, pricing, context limits, and availability after the announced shutdown date.
Bottom line
GPT-Realtime is a specialized OpenAI model for fast speech-to-speech and multimodal sessions. It combines text and audio input and output, image input, persistent realtime connections, and function calling in a single model-oriented workflow. Those capabilities make it a practical fit for voice agents and conversational applications, while its pricing, limited web and structured-output support, modest editorial reasoning and coding ratings, and scheduled API shutdown limit its appeal for new long-term systems.

