What GPT-4o Realtime was
GPT-4o Realtime was OpenAI’s preview realtime model for interactive text-and-audio applications. Its canonical API identifier was gpt-4o-realtime-preview. OpenAI introduced it with the Realtime API on October 1, 2024, positioning it for conversations where waiting for a separate transcription, text-generation, and speech-synthesis pipeline would add noticeable delay.
Unlike a conventional text-only language-model endpoint, the model could receive audio directly and produce audio directly. That allowed an application to maintain a conversational session in which a user could speak, be interrupted, and receive a spoken response. The Realtime API exposed this interaction through WebRTC or WebSocket connections.
The most important qualification is its current status: GPT-4o Realtime is retired. OpenAI announced its deprecation in September 2025 and removed the model family from the API on May 7, 2026. The model therefore should not be selected for a new production integration, even though its historical specifications remain useful for understanding OpenAI’s earlier realtime architecture.
Modalities and core capabilities
GPT-4o Realtime accepted both text and audio input and could return both text and audio output. In practical terms, developers could use it for spoken conversations, text fallback interfaces, or applications that combined typed instructions with voice responses.
- Input: text and audio.
- Output: text and audio, including direct speech output.
- Connection methods: WebRTC and WebSocket realtime sessions.
- Tool support: function calling for triggering application actions or retrieving external information.
- Input caching: cached text and audio input pricing was supported.
The model did not support image or video input, image or video generation, native web search, fine-tuning, or structured outputs according to the supplied model documentation. Its audio capability should therefore not be confused with general visual multimodality: GPT-4o Realtime was a text-and-audio model, not an image-and-video system.
Why low latency mattered
Voice interaction is sensitive to pauses. A traditional voice assistant may first transcribe speech, send the transcript to a language model, and then pass the answer to a speech synthesizer. Each handoff can add delay and can also discard information contained in the original audio, such as timing or conversational cues. GPT-4o Realtime was designed to handle text and speech within a realtime session instead.
This made the model a natural fit for assistants that needed to respond while a conversation was in progress. Realtime sessions also supported interruption handling through the API, which is important in voice interfaces: a user can begin speaking before a long answer has finished rather than waiting for the system to complete its turn.
These strengths were architectural rather than a guarantee that every application would have identical latency. Network conditions, client implementation, session management, and the surrounding application all affected the final user experience.
Technical specifications and pricing
The following values come from the documented GPT-4o Realtime model specifications. Token prices differed between text and audio, and audio output was substantially more expensive than text output.
| Specification | Documented value |
|---|---|
| Canonical model ID | gpt-4o-realtime-preview |
| Context window | 32,000 tokens |
| Maximum output | 4,096 tokens |
| Knowledge cutoff | October 1, 2023 |
| Text input | $5 per 1 million tokens |
| Cached text input | $2.50 per 1 million tokens |
| Text output | $20 per 1 million tokens |
| Audio input | $40 per 1 million tokens |
| Cached audio input | $2.50 per 1 million tokens |
| Audio output | $80 per 1 million tokens |
The 32,000-token context window limited how much conversation history and application context could be retained in one session. Long-running assistants therefore needed to manage history carefully, especially when audio exchanges consumed substantial token volume. The 4,096-token maximum output was the documented output ceiling, although a normal spoken response would typically be much shorter than that limit.
Tools, reasoning, and coding considerations
GPT-4o Realtime supported function calling. A function call lets the model request that the host application perform a defined operation, such as looking up account information, changing a reservation, or retrieving data from an internal service. The application—not the model—remained responsible for executing the function and returning its result.
Function calling made the model more useful than a voice-only answer generator because a conversation could lead to an actual application action. However, the supplied research does not document native web search, structured JSON output, or fine-tuning support. Applications that require strict machine-readable output should not assume that a spoken or text response can replace a structured-output-capable model.
The supplied evaluation record assigns GPT-4o Realtime a reasoning score of 6 and a coding score of 7. These are editorial or catalog evaluations, not scores published as official OpenAI benchmarks, so they should be treated as relative guidance rather than verified performance measurements. The model’s primary distinction was realtime speech interaction, not specialized reasoning or software-development performance.
Speed, cost, and capability trade-offs
GPT-4o Realtime traded higher audio costs for direct, low-latency interaction. Its documented text prices were lower than its audio prices, while audio output cost $80 per 1 million audio tokens. That pricing structure made voice design and response length important parts of cost control. Caching could reduce the price of repeated text and audio input, but it did not remove the separate cost of generated audio.
The supplied catalog gives the model a speed score of 9 and a cost score of 4. Those values are editorial assessments, not provider-published measurements. They summarize a practical trade-off: the model was attractive when conversational responsiveness mattered more than minimizing inference cost, but less attractive for large volumes of inexpensive text generation.
Its multimodal input and output fields are both marked as supported because it handled audio as well as text. Its streaming field is marked unsupported in the model-page feature record, even though the Realtime API provided streamed session interactions. This distinction reflects the supplied catalog’s exact feature field and should not be read as evidence that realtime sessions were non-streaming.
Best use cases before retirement
Before its API shutdown, GPT-4o Realtime was particularly appropriate for:
- Voice assistants: spoken help with rapid turn-taking and interruption handling.
- Language learning: conversational practice with text or audio interaction.
- Live translation: low-latency spoken exchanges where waiting for a full transcript would be inconvenient.
- Customer support: voice agents that could call application functions to retrieve information or initiate supported actions.
- Accessibility tools: interfaces that relied on speech rather than typing or reading a long text response.
These use cases benefited from the model’s combination of audio input, audio output, and function calling. They also required careful handling of permissions, application-side tool execution, session state, and usage costs.
When not to choose GPT-4o Realtime
GPT-4o Realtime is not appropriate for a new deployment because API access has ended. OpenAI recommended gpt-realtime-1.5 as its replacement, and current projects should evaluate that model family instead of building against the retired preview identifier.
Even before retirement, another type of model would have been more suitable when the application needed image or video processing, strict structured output, fine-tuning, native web search, or batch processing. A text-oriented model could also be more economical when voice interaction was unnecessary. Conversely, a separate transcription-and-text-to-speech architecture might offer more component-level control, even though it could introduce additional latency and integration complexity.
The decision was therefore not simply about whether GPT-4o Realtime could produce a good answer. It depended on whether direct speech-to-speech interaction justified audio pricing and the model’s limits. For a fast spoken conversation, that trade-off was its central value. For new systems today, its retired status outweighs those historical advantages.
Bottom line
GPT-4o Realtime was an important OpenAI preview model for low-latency text-and-audio conversations. It combined direct audio interaction, realtime WebRTC or WebSocket sessions, interruption-aware conversation handling, and function calling in a single model experience. Its limitations included a 32,000-token context window, 4,096-token maximum output, relatively high audio pricing, no image or video support, no structured outputs, and no fine-tuning.
Those specifications explain why it was useful for voice assistants and interactive support applications, but its retirement on May 7, 2026 determines its practical recommendation now: treat GPT-4o Realtime as a historical model record, and use OpenAI’s recommended current realtime replacement for new work.

