GPT-Realtime

GPT-Realtime-1.5

by OpenAI · Active and currently available

OpenAI’s GPT-Realtime-1.5 is a fast, non-reasoning speech-to-speech model for realtime voice agents. It supports text, image, and audio input, text and audio output, function calling, a 32,000-token context window, and modality-specific pricing. Its main trade-offs are higher audio costs, no structured outputs or fine-tuning, and limited suitability for complex reasoning.

Text Speech Reasoning Coding
GPT-Realtime-1.5 is OpenAI’s production-focused model for realtime voice applications. Through the Realtime API, it can listen to audio, respond with synthesized speech, use visual context, and call application tools during an ongoing conversation. It is optimized for fast, natural interaction rather than deep multi-step reasoning, making its main trade-off straightforward: lower-latency voice experiences in exchange for less suitability for complex analytical work.
Outputs

What GPT-Realtime-1.5 can produce

Text Speech
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Tool use Prompt caching Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
4/10 Coding
9/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime
Model type Realtime Audio
Context window 32K tokens
Maximum output 4K tokens
Knowledge cutoff 2024-09-30
Release date 2026-02-23
Status Active and currently available
Knowledge cutoff notes

The official model reference lists September 30, 2024 as the model's knowledge cutoff. External tools or retrieval can provide newer information during use but do not change the underlying cutoff.

Model notes

GPT-Realtime-1.5 is a fast, reliable non-reasoning speech-to-speech model. It accepts text, images, and audio, and produces text or audio. The model reference lists conventional token streaming as unsupported, although realtime sessions stream audio and event data over WebRTC or WebSocket connections. Function calling is supported. Structured outputs and fine-tuning are not supported. OpenAI lists a September 30, 2024 knowledge cutoff. No exact deprecation or shutdown date is published for this model as of September 23, 2026. Older gpt-4o-realtime-preview models were scheduled for removal on May 7, 2026 with GPT-Realtime-1.5 as a recommended replacement.

Cost

Model pricing

Input $4.00 per 1M text tokens; $32.00 per 1M audio tokens; $5.00 per 1M image tokens. Cached input: $0.40 per 1M text or audio tokens and $0.50 per 1M image tokens.
Output $16.00 per 1M text tokens; $64.00 per 1M audio tokens.
Model guide

GPT-Realtime-1.5: Features, Pricing, Context and Audio Support

GPT-Realtime-1.5 is OpenAI’s non-reasoning realtime audio model for building low-latency speech-to-speech applications. It accepts text, images, and audio, produces text or synthesized audio, and supports function calling for voice agents, customer-service systems, and interactive assistants.

What is GPT-Realtime-1.5?

GPT-Realtime-1.5 is OpenAI’s realtime audio model for applications that need to hear a user and respond naturally while the conversation is still in progress. Its primary design is speech-to-speech: an application can send spoken audio and receive spoken audio without separately connecting a speech-recognition model, a text language model, and a text-to-speech service.

OpenAI released the model to the Realtime API on February 23, 2026. The provider describes it as a fast, reliable, non-reasoning realtime model with improved instruction following, tool calling, multilingual handling, and voice quality compared with earlier realtime preview models. Those are provider positioning claims; the specifications and pricing below are the concrete documented details supplied for this model.

GPT-Realtime-1.5 sits in OpenAI’s realtime model lineup rather than serving as a general-purpose replacement for every text or reasoning model. It is intended for voice agents, customer-support systems, interactive assistants, and other applications where response timing and conversational continuity matter more than extended internal analysis.

Supported modalities and capabilities

The model accepts text, images, and audio as input. It produces text or audio as output. Images provide visual context but are not generated by the model, and video input is not supported.

  • Text input and output: Supported.
  • Audio input and output: Supported for realtime voice conversations.
  • Image input: Supported for visual context.
  • Image and video output: Not supported.
  • Video input: Not supported.
  • Function calling: Supported, allowing the model to request actions from application tools or backend services.
  • Structured outputs: Not supported according to the model reference.
  • Fine-tuning: Not supported.

Function calling is particularly important for voice agents. For example, a scheduling assistant could collect information by voice, call a calendar or booking system, and then tell the user the result. The model does not independently perform those backend actions; the application must define the tools and execute approved requests.

Context window and technical specifications

SpecificationGPT-Realtime-1.5
Model IDgpt-realtime-1.5
Model typeRealtime audio and speech-to-speech
Context window32,000 tokens
Maximum output4,096 tokens
Knowledge cutoffSeptember 30, 2024
Fine-tuningNot supported
Structured outputsNot supported
Function callingSupported

A token is a unit used to represent text and other model input or output data. The 32,000-token context window limits how much conversation history, instructions, tool information, and multimodal context can be kept available at one time. The 4,096-token maximum applies to model output; realtime audio sessions also emit incremental events as the response is produced.

GPT-Realtime-1.5 is accessed through realtime interfaces, including WebRTC and WebSocket-based workflows. Audio can be sent in chunks, while the service emits audio and transcript events during a response. The model reference lists conventional token streaming as unsupported, so this should not be confused with a standard text-completion streaming setting. Realtime event transport and incremental audio delivery are still central to its intended use.

GPT-Realtime-1.5 pricing

Pricing is based on modality-specific token usage rather than a single flat request price.

Usage typePrice per 1 million tokens
Text input$4.00
Cached text input$0.40
Text output$16.00
Audio input$32.00
Cached audio input$0.40
Audio output$64.00
Image input$5.00
Cached image input$0.50

Audio is substantially more expensive than text on an uncached basis, especially for output. This means a voice application should estimate the amount of listening and speaking it will generate, rather than budgeting only from the number of conversations or API requests. Cached input pricing can reduce the cost of repeated audio, text, or image context when the relevant input is eligible for caching.

The model can still be economically attractive for applications where direct speech-to-speech interaction reduces the engineering and latency costs of chaining separate transcription, language, and speech-generation services. Whether it is cheaper overall depends on conversation length, audio volume, caching, tool usage, and the alternative architecture.

Reasoning, coding, and tool use

GPT-Realtime-1.5 is explicitly a non-reasoning model. It can follow instructions, maintain a live conversation, and select or request application tools, but it is not the best choice for difficult multi-step planning, extensive analysis, or tasks that demand the strongest reasoning reliability. A voice interface may use another service for complex reasoning while retaining GPT-Realtime-1.5 for conversational input and output, although that introduces additional system complexity and latency.

The supplied research does not identify a dedicated coding capability or coding-specific optimization for GPT-Realtime-1.5. It can process text and call tools, so an application may expose coding or software-development actions to it, but that should not be interpreted as evidence that it is a specialist coding model. For coding-heavy work, a general coding-oriented model may be more appropriate.

Function calling is a verified capability. In practice, the application defines functions such as search_customer, book_appointment, or check_order_status. GPT-Realtime-1.5 can determine when a function is needed and provide arguments, while the application validates those arguments, performs the operation, and returns the result. Confirmation rules are important for actions involving payments, cancellations, account changes, or other irreversible effects.

Main strengths and limitations

Where GPT-Realtime-1.5 is strongest

  • Low-latency voice interaction: It is designed for conversations where waiting for separate transcription and speech-generation stages would make the experience feel slow.
  • Direct speech-to-speech operation: Audio can remain the main interaction format instead of being converted through a visibly separate pipeline.
  • Tool-connected assistants: Function calling allows a voice agent to retrieve information or initiate application actions.
  • Multimodal context: Text, audio, and images can be combined as input, which is useful for assistants that need to discuss a document, screen, or image while speaking.
  • Realtime transport: WebRTC and WebSocket workflows support ongoing sessions and incremental audio or event delivery.

Important limitations

  • It is not a reasoning model and may be a weaker fit for complex planning or difficult analytical tasks.
  • Structured outputs are not supported, which makes it less suitable for workflows that require guaranteed schema-conforming JSON directly from the model.
  • Fine-tuning is not supported.
  • It does not generate images or video and does not accept video input.
  • Audio input and output pricing is higher than text pricing.
  • The September 30, 2024 knowledge cutoff means current information must come from application tools or retrieval systems.
  • Realtime voice systems require careful testing of interruptions, pronunciation, identity details, confirmations, and fallback behavior.

These limitations are not merely catalog details. For example, a customer-support agent can use a live order-status tool, but the model’s built-in knowledge does not automatically become current. Similarly, an application that requires strict machine-readable output may need a separate processing step because structured outputs are not supported.

Best use cases

GPT-Realtime-1.5 is a strong candidate when the user’s main interaction is spoken and the application needs an immediate conversational response. Suitable examples include:

  • Customer-support and contact-center voice agents.
  • Sales, scheduling, and service assistants that connect to business tools.
  • Interactive voice-response systems with backend actions.
  • Language-learning and conversational-practice applications.
  • Voice interfaces for browsers, mobile applications, telephony, and SIP systems.
  • Assistants that need to discuss images or visual context while responding aloud.

It is less suitable for offline batch processing, image or video generation, fine-tuned domain models, strict structured-data extraction, or workloads dominated by complex reasoning. A text-first model may be more economical when users do not need audio, while a reasoning-oriented model may be more appropriate when correctness depends on long chains of analysis rather than rapid conversation.

When to choose GPT-Realtime-1.5

Choose GPT-Realtime-1.5 when the core product experience is a responsive voice conversation and the system benefits from built-in audio input and output. It is especially appropriate when reducing integration steps, handling interruptions, and calling tools during a live session are more important than maximizing reasoning depth.

Choose another type of model when the workload is primarily written, highly analytical, coding-intensive, or dependent on strict JSON schemas. Text models can offer a better cost profile for text-only traffic, and reasoning or specialist models can be preferable for difficult planning and technical analysis. GPT-Realtime-1.5 can still serve as the voice layer in a larger system, but combining it with another model adds orchestration work and may reduce the simplicity advantage of a direct speech-to-speech design.

Availability and current status

As of September 23, 2026, GPT-Realtime-1.5 is listed in OpenAI’s current model catalog and remains available through the Realtime API. OpenAI’s documented retirement guidance for older GPT-4o realtime preview models recommends GPT-Realtime-1.5 as a replacement. No deprecation or shutdown date is published for GPT-Realtime-1.5 in the supplied research.

Before production deployment, developers should verify current pricing, model availability, transport requirements, and retirement notices in OpenAI’s documentation. They should also measure real conversation costs using representative audio durations, because audio output and input usage can dominate the bill even when the number of sessions appears modest.


Answers to Frequently Asked Questions

Is GPT-Realtime-1.5 suitable for complex reasoning or structured data extraction?
GPT-Realtime-1.5 is explicitly a non-reasoning model, so it is better suited to fast conversational interactions and tool-connected voice agents than to difficult planning, extensive analysis, or reasoning-intensive tasks. Structured outputs are not supported, so workflows requiring guaranteed schema-conforming JSON may need a separate processing step or another model.
What are GPT-Realtime-1.5’s context window and output limits?
GPT-Realtime-1.5 has a 32,000-token context window and a maximum output of 4,096 tokens. The context window includes conversation history, instructions, tool information, and multimodal context. Realtime sessions can still deliver incremental audio and event updates as responses are produced.
How much does GPT-Realtime-1.5 cost?
GPT-Realtime-1.5 pricing is charged per 1 million modality-specific tokens: $4.00 for text input, $0.40 for cached text input, $16.00 for text output, $32.00 for audio input, $0.40 for cached audio input, $64.00 for audio output, $5.00 for image input, and $0.50 for cached image input. Audio usage is substantially more expensive than uncached text usage.
What is GPT-Realtime-1.5 designed for?
GPT-Realtime-1.5 is OpenAI’s realtime speech-to-speech model for applications that need to hear users and respond naturally during an ongoing conversation. It is designed for voice agents, customer-support systems, interactive assistants, scheduling tools, and other low-latency conversational applications.
What modalities and capabilities does GPT-Realtime-1.5 support?
The model accepts text, audio, and images as input and produces text or audio as output. It supports realtime voice conversations, image-based context, and function calling for connecting applications to backend tools. It does not support video input, image or video generation, structured outputs, or fine-tuning.


Sources 7
Provider

About OpenAI