GPT-4o

GPT-4o Realtime

by OpenAI · Retired; API access ended 2026-05-07

GPT-4o Realtime was OpenAI’s preview speech-to-speech model for low-latency conversations over WebRTC and WebSocket. It supported text and audio input and output, function calling, and cached inputs, but had high audio costs and was retired from the API on May 7, 2026.

Text Speech Reasoning Coding
GPT-4o Realtime was built for applications that needed natural, fast voice interaction without assembling separate speech-recognition, language-model, and text-to-speech systems. It supported text and audio input and output, function calling, persistent realtime sessions, and low-latency communication over WebRTC or WebSocket. Although it was well suited to voice assistants, tutoring, translation, and customer support, its preview status, audio pricing, 32,000-token context window, and lack of structured outputs limited some use cases. OpenAI retired the model from the API on May 7, 2026 and recommended migrating to gpt-realtime-1.5.
Outputs

What GPT-4o Realtime can produce

Text Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Tool use Prompt caching Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
7/10 Coding
9/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o
Model type Multimodal
Context window 32K tokens
Maximum output 4K tokens
Knowledge cutoff 2023-10-01
Release date 2024-10-01
Status Retired; API access ended 2026-05-07
Deprecation date 2025-09-15
Shutdown date 2026-05-07
Knowledge cutoff notes

The official GPT-4o Realtime model page lists October 1, 2023 as the model's knowledge cutoff.

Model notes

Canonical model identifier was gpt-4o-realtime-preview. OpenAI described it as a preview release capable of realtime text and audio interaction over WebRTC or WebSocket. It supported function calling and cached text and audio input. The model had a 32,000-token context window, 4,096-token maximum output, and an October 1, 2023 knowledge cutoff. OpenAI announced deprecation in September 2025 and removed the model from the API on May 7, 2026, recommending gpt-realtime-1.5 as the replacement. The model page's streaming field is marked unsupported even though the Realtime API provided streamed session interactions; this record follows the exact model-page feature field.

Cost

Model pricing

Input $5 per 1M text tokens; $40 per 1M audio tokens; cached text and audio input $2.50 per 1M tokens
Output $20 per 1M text tokens; $80 per 1M audio tokens
Model guide

GPT-4o Realtime: OpenAI’s Retired Speech-to-Speech Model

GPT-4o Realtime was OpenAI’s preview model for low-latency text-and-audio conversations. It accepted text or audio and returned text or audio through realtime WebRTC or WebSocket sessions, with function calling and cached input support. The model was retired from the API on May 7, 2026, so it is now primarily useful as historical documentation rather than as a choice for new deployments.

What GPT-4o Realtime was

GPT-4o Realtime was OpenAI’s preview realtime model for interactive text-and-audio applications. Its canonical API identifier was gpt-4o-realtime-preview. OpenAI introduced it with the Realtime API on October 1, 2024, positioning it for conversations where waiting for a separate transcription, text-generation, and speech-synthesis pipeline would add noticeable delay.

Unlike a conventional text-only language-model endpoint, the model could receive audio directly and produce audio directly. That allowed an application to maintain a conversational session in which a user could speak, be interrupted, and receive a spoken response. The Realtime API exposed this interaction through WebRTC or WebSocket connections.

The most important qualification is its current status: GPT-4o Realtime is retired. OpenAI announced its deprecation in September 2025 and removed the model family from the API on May 7, 2026. The model therefore should not be selected for a new production integration, even though its historical specifications remain useful for understanding OpenAI’s earlier realtime architecture.

Modalities and core capabilities

GPT-4o Realtime accepted both text and audio input and could return both text and audio output. In practical terms, developers could use it for spoken conversations, text fallback interfaces, or applications that combined typed instructions with voice responses.

  • Input: text and audio.
  • Output: text and audio, including direct speech output.
  • Connection methods: WebRTC and WebSocket realtime sessions.
  • Tool support: function calling for triggering application actions or retrieving external information.
  • Input caching: cached text and audio input pricing was supported.

The model did not support image or video input, image or video generation, native web search, fine-tuning, or structured outputs according to the supplied model documentation. Its audio capability should therefore not be confused with general visual multimodality: GPT-4o Realtime was a text-and-audio model, not an image-and-video system.

Why low latency mattered

Voice interaction is sensitive to pauses. A traditional voice assistant may first transcribe speech, send the transcript to a language model, and then pass the answer to a speech synthesizer. Each handoff can add delay and can also discard information contained in the original audio, such as timing or conversational cues. GPT-4o Realtime was designed to handle text and speech within a realtime session instead.

This made the model a natural fit for assistants that needed to respond while a conversation was in progress. Realtime sessions also supported interruption handling through the API, which is important in voice interfaces: a user can begin speaking before a long answer has finished rather than waiting for the system to complete its turn.

These strengths were architectural rather than a guarantee that every application would have identical latency. Network conditions, client implementation, session management, and the surrounding application all affected the final user experience.

Technical specifications and pricing

The following values come from the documented GPT-4o Realtime model specifications. Token prices differed between text and audio, and audio output was substantially more expensive than text output.

SpecificationDocumented value
Canonical model IDgpt-4o-realtime-preview
Context window32,000 tokens
Maximum output4,096 tokens
Knowledge cutoffOctober 1, 2023
Text input$5 per 1 million tokens
Cached text input$2.50 per 1 million tokens
Text output$20 per 1 million tokens
Audio input$40 per 1 million tokens
Cached audio input$2.50 per 1 million tokens
Audio output$80 per 1 million tokens

The 32,000-token context window limited how much conversation history and application context could be retained in one session. Long-running assistants therefore needed to manage history carefully, especially when audio exchanges consumed substantial token volume. The 4,096-token maximum output was the documented output ceiling, although a normal spoken response would typically be much shorter than that limit.

Tools, reasoning, and coding considerations

GPT-4o Realtime supported function calling. A function call lets the model request that the host application perform a defined operation, such as looking up account information, changing a reservation, or retrieving data from an internal service. The application—not the model—remained responsible for executing the function and returning its result.

Function calling made the model more useful than a voice-only answer generator because a conversation could lead to an actual application action. However, the supplied research does not document native web search, structured JSON output, or fine-tuning support. Applications that require strict machine-readable output should not assume that a spoken or text response can replace a structured-output-capable model.

The supplied evaluation record assigns GPT-4o Realtime a reasoning score of 6 and a coding score of 7. These are editorial or catalog evaluations, not scores published as official OpenAI benchmarks, so they should be treated as relative guidance rather than verified performance measurements. The model’s primary distinction was realtime speech interaction, not specialized reasoning or software-development performance.

Speed, cost, and capability trade-offs

GPT-4o Realtime traded higher audio costs for direct, low-latency interaction. Its documented text prices were lower than its audio prices, while audio output cost $80 per 1 million audio tokens. That pricing structure made voice design and response length important parts of cost control. Caching could reduce the price of repeated text and audio input, but it did not remove the separate cost of generated audio.

The supplied catalog gives the model a speed score of 9 and a cost score of 4. Those values are editorial assessments, not provider-published measurements. They summarize a practical trade-off: the model was attractive when conversational responsiveness mattered more than minimizing inference cost, but less attractive for large volumes of inexpensive text generation.

Its multimodal input and output fields are both marked as supported because it handled audio as well as text. Its streaming field is marked unsupported in the model-page feature record, even though the Realtime API provided streamed session interactions. This distinction reflects the supplied catalog’s exact feature field and should not be read as evidence that realtime sessions were non-streaming.

Best use cases before retirement

Before its API shutdown, GPT-4o Realtime was particularly appropriate for:

  • Voice assistants: spoken help with rapid turn-taking and interruption handling.
  • Language learning: conversational practice with text or audio interaction.
  • Live translation: low-latency spoken exchanges where waiting for a full transcript would be inconvenient.
  • Customer support: voice agents that could call application functions to retrieve information or initiate supported actions.
  • Accessibility tools: interfaces that relied on speech rather than typing or reading a long text response.

These use cases benefited from the model’s combination of audio input, audio output, and function calling. They also required careful handling of permissions, application-side tool execution, session state, and usage costs.

When not to choose GPT-4o Realtime

GPT-4o Realtime is not appropriate for a new deployment because API access has ended. OpenAI recommended gpt-realtime-1.5 as its replacement, and current projects should evaluate that model family instead of building against the retired preview identifier.

Even before retirement, another type of model would have been more suitable when the application needed image or video processing, strict structured output, fine-tuning, native web search, or batch processing. A text-oriented model could also be more economical when voice interaction was unnecessary. Conversely, a separate transcription-and-text-to-speech architecture might offer more component-level control, even though it could introduce additional latency and integration complexity.

The decision was therefore not simply about whether GPT-4o Realtime could produce a good answer. It depended on whether direct speech-to-speech interaction justified audio pricing and the model’s limits. For a fast spoken conversation, that trade-off was its central value. For new systems today, its retired status outweighs those historical advantages.

Bottom line

GPT-4o Realtime was an important OpenAI preview model for low-latency text-and-audio conversations. It combined direct audio interaction, realtime WebRTC or WebSocket sessions, interruption-aware conversation handling, and function calling in a single model experience. Its limitations included a 32,000-token context window, 4,096-token maximum output, relatively high audio pricing, no image or video support, no structured outputs, and no fine-tuning.

Those specifications explain why it was useful for voice assistants and interactive support applications, but its retirement on May 7, 2026 determines its practical recommendation now: treat GPT-4o Realtime as a historical model record, and use OpenAI’s recommended current realtime replacement for new work.


Answers to Frequently Asked Questions

What replaced GPT-4o Realtime?
OpenAI recommended gpt-realtime-1.5 as the replacement for GPT-4o Realtime. New projects should evaluate the current realtime model family rather than using the retired gpt-4o-realtime-preview identifier.
What could GPT-4o Realtime be used for?
Before its retirement, GPT-4o Realtime was suited to voice assistants, language-learning applications, live translation, customer-support agents, and accessibility tools. Its main advantages were direct audio interaction, low-latency turn-taking, interruption handling, and function calling.
Is GPT-4o Realtime still available?
No. OpenAI deprecated GPT-4o Realtime in September 2025 and removed the model family from the API on May 7, 2026. It should not be used for new production integrations.
What was the model ID for GPT-4o Realtime?
The canonical API identifier was gpt-4o-realtime-preview.
What was GPT-4o Realtime?
GPT-4o Realtime was OpenAI’s preview model for low-latency text-and-audio conversations. It accepted text and audio input, returned text and audio output, and supported realtime sessions through WebRTC or WebSocket connections.


Sources 3
Provider

About OpenAI