What is GPT-Realtime-1.5?
GPT-Realtime-1.5 is OpenAI’s realtime audio model for applications that need to hear a user and respond naturally while the conversation is still in progress. Its primary design is speech-to-speech: an application can send spoken audio and receive spoken audio without separately connecting a speech-recognition model, a text language model, and a text-to-speech service.
OpenAI released the model to the Realtime API on February 23, 2026. The provider describes it as a fast, reliable, non-reasoning realtime model with improved instruction following, tool calling, multilingual handling, and voice quality compared with earlier realtime preview models. Those are provider positioning claims; the specifications and pricing below are the concrete documented details supplied for this model.
GPT-Realtime-1.5 sits in OpenAI’s realtime model lineup rather than serving as a general-purpose replacement for every text or reasoning model. It is intended for voice agents, customer-support systems, interactive assistants, and other applications where response timing and conversational continuity matter more than extended internal analysis.
Supported modalities and capabilities
The model accepts text, images, and audio as input. It produces text or audio as output. Images provide visual context but are not generated by the model, and video input is not supported.
- Text input and output: Supported.
- Audio input and output: Supported for realtime voice conversations.
- Image input: Supported for visual context.
- Image and video output: Not supported.
- Video input: Not supported.
- Function calling: Supported, allowing the model to request actions from application tools or backend services.
- Structured outputs: Not supported according to the model reference.
- Fine-tuning: Not supported.
Function calling is particularly important for voice agents. For example, a scheduling assistant could collect information by voice, call a calendar or booking system, and then tell the user the result. The model does not independently perform those backend actions; the application must define the tools and execute approved requests.
Context window and technical specifications
| Specification | GPT-Realtime-1.5 |
|---|---|
| Model ID | gpt-realtime-1.5 |
| Model type | Realtime audio and speech-to-speech |
| Context window | 32,000 tokens |
| Maximum output | 4,096 tokens |
| Knowledge cutoff | September 30, 2024 |
| Fine-tuning | Not supported |
| Structured outputs | Not supported |
| Function calling | Supported |
A token is a unit used to represent text and other model input or output data. The 32,000-token context window limits how much conversation history, instructions, tool information, and multimodal context can be kept available at one time. The 4,096-token maximum applies to model output; realtime audio sessions also emit incremental events as the response is produced.
GPT-Realtime-1.5 is accessed through realtime interfaces, including WebRTC and WebSocket-based workflows. Audio can be sent in chunks, while the service emits audio and transcript events during a response. The model reference lists conventional token streaming as unsupported, so this should not be confused with a standard text-completion streaming setting. Realtime event transport and incremental audio delivery are still central to its intended use.
GPT-Realtime-1.5 pricing
Pricing is based on modality-specific token usage rather than a single flat request price.
| Usage type | Price per 1 million tokens |
|---|---|
| Text input | $4.00 |
| Cached text input | $0.40 |
| Text output | $16.00 |
| Audio input | $32.00 |
| Cached audio input | $0.40 |
| Audio output | $64.00 |
| Image input | $5.00 |
| Cached image input | $0.50 |
Audio is substantially more expensive than text on an uncached basis, especially for output. This means a voice application should estimate the amount of listening and speaking it will generate, rather than budgeting only from the number of conversations or API requests. Cached input pricing can reduce the cost of repeated audio, text, or image context when the relevant input is eligible for caching.
The model can still be economically attractive for applications where direct speech-to-speech interaction reduces the engineering and latency costs of chaining separate transcription, language, and speech-generation services. Whether it is cheaper overall depends on conversation length, audio volume, caching, tool usage, and the alternative architecture.
Reasoning, coding, and tool use
GPT-Realtime-1.5 is explicitly a non-reasoning model. It can follow instructions, maintain a live conversation, and select or request application tools, but it is not the best choice for difficult multi-step planning, extensive analysis, or tasks that demand the strongest reasoning reliability. A voice interface may use another service for complex reasoning while retaining GPT-Realtime-1.5 for conversational input and output, although that introduces additional system complexity and latency.
The supplied research does not identify a dedicated coding capability or coding-specific optimization for GPT-Realtime-1.5. It can process text and call tools, so an application may expose coding or software-development actions to it, but that should not be interpreted as evidence that it is a specialist coding model. For coding-heavy work, a general coding-oriented model may be more appropriate.
Function calling is a verified capability. In practice, the application defines functions such as search_customer, book_appointment, or check_order_status. GPT-Realtime-1.5 can determine when a function is needed and provide arguments, while the application validates those arguments, performs the operation, and returns the result. Confirmation rules are important for actions involving payments, cancellations, account changes, or other irreversible effects.
Main strengths and limitations
Where GPT-Realtime-1.5 is strongest
- Low-latency voice interaction: It is designed for conversations where waiting for separate transcription and speech-generation stages would make the experience feel slow.
- Direct speech-to-speech operation: Audio can remain the main interaction format instead of being converted through a visibly separate pipeline.
- Tool-connected assistants: Function calling allows a voice agent to retrieve information or initiate application actions.
- Multimodal context: Text, audio, and images can be combined as input, which is useful for assistants that need to discuss a document, screen, or image while speaking.
- Realtime transport: WebRTC and WebSocket workflows support ongoing sessions and incremental audio or event delivery.
Important limitations
- It is not a reasoning model and may be a weaker fit for complex planning or difficult analytical tasks.
- Structured outputs are not supported, which makes it less suitable for workflows that require guaranteed schema-conforming JSON directly from the model.
- Fine-tuning is not supported.
- It does not generate images or video and does not accept video input.
- Audio input and output pricing is higher than text pricing.
- The September 30, 2024 knowledge cutoff means current information must come from application tools or retrieval systems.
- Realtime voice systems require careful testing of interruptions, pronunciation, identity details, confirmations, and fallback behavior.
These limitations are not merely catalog details. For example, a customer-support agent can use a live order-status tool, but the model’s built-in knowledge does not automatically become current. Similarly, an application that requires strict machine-readable output may need a separate processing step because structured outputs are not supported.
Best use cases
GPT-Realtime-1.5 is a strong candidate when the user’s main interaction is spoken and the application needs an immediate conversational response. Suitable examples include:
- Customer-support and contact-center voice agents.
- Sales, scheduling, and service assistants that connect to business tools.
- Interactive voice-response systems with backend actions.
- Language-learning and conversational-practice applications.
- Voice interfaces for browsers, mobile applications, telephony, and SIP systems.
- Assistants that need to discuss images or visual context while responding aloud.
It is less suitable for offline batch processing, image or video generation, fine-tuned domain models, strict structured-data extraction, or workloads dominated by complex reasoning. A text-first model may be more economical when users do not need audio, while a reasoning-oriented model may be more appropriate when correctness depends on long chains of analysis rather than rapid conversation.
When to choose GPT-Realtime-1.5
Choose GPT-Realtime-1.5 when the core product experience is a responsive voice conversation and the system benefits from built-in audio input and output. It is especially appropriate when reducing integration steps, handling interruptions, and calling tools during a live session are more important than maximizing reasoning depth.
Choose another type of model when the workload is primarily written, highly analytical, coding-intensive, or dependent on strict JSON schemas. Text models can offer a better cost profile for text-only traffic, and reasoning or specialist models can be preferable for difficult planning and technical analysis. GPT-Realtime-1.5 can still serve as the voice layer in a larger system, but combining it with another model adds orchestration work and may reduce the simplicity advantage of a direct speech-to-speech design.
Availability and current status
As of September 23, 2026, GPT-Realtime-1.5 is listed in OpenAI’s current model catalog and remains available through the Realtime API. OpenAI’s documented retirement guidance for older GPT-4o realtime preview models recommends GPT-Realtime-1.5 as a replacement. No deprecation or shutdown date is published for GPT-Realtime-1.5 in the supplied research.
Before production deployment, developers should verify current pricing, model availability, transport requirements, and retirement notices in OpenAI’s documentation. They should also measure real conversation costs using representative audio durations, because audio output and input usage can dominate the bill even when the number of sessions appears modest.

