GPT-Realtime

GPT-Realtime-Whisper

by OpenAI · Current and available through the OpenAI Realtime API for realtime transcription

OpenAI’s GPT-Realtime-Whisper provides low-latency streaming speech-to-text for live audio. It returns realtime transcript updates, costs $0.017 per minute of audio, has a 16,000-token context window and a 2,000-token maximum output. The model is designed for captions, meetings, call analysis and voice-agent input, but does not generate speech or support function calling and structured outputs.

Text Reasoning Coding
GPT-Realtime-Whisper is a specialized OpenAI model for real-time transcription rather than a general conversational assistant. It accepts live audio and produces streaming text, allowing applications to display or process transcript updates before an audio session ends. Its main trade-off is deliberate specialization: it is designed for speed and continuous speech-to-text, not spoken responses, general reasoning, coding, function calling or structured data generation.
Outputs

What GPT-Realtime-Whisper can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Batch API
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime
Model type Other
Context window 16K tokens
Maximum output 2K tokens
Knowledge cutoff September 30, 2024
Release date May 7, 2026
Status Current and available through the OpenAI Realtime API for realtime transcription
Knowledge cutoff notes

The official model page specifies September 30, 2024 as the model's knowledge cutoff. This cutoff applies to the underlying model and is not changed by realtime audio input or surrounding application context.

Model notes

GPT-Realtime-Whisper is a specialist transcription model rather than a conversational reasoning model. It produces streaming text from live audio and is priced by audio duration. The official model page lists text input and output, audio input only, a 16,000-token context window, and a 2,000-token maximum output. Function calling, structured outputs, and fine-tuning are listed as unsupported. The editorial reasoning and coding scores are intentionally low because those capabilities are outside the model's primary purpose.

Cost

Model pricing

Input $0.017 per minute of audio
Output Included in the audio-duration transcription price; no separate text-output price documented
Model guide

GPT-Realtime-Whisper: Pricing, Capabilities, Limits and API Support

GPT-Realtime-Whisper is OpenAI’s low-latency streaming speech-to-text model. It converts live audio into incremental text updates for captions, meeting notes, call analysis, voice-agent input, accessibility workflows and other applications that need a transcript while someone is still speaking.

What is GPT-Realtime-Whisper?

GPT-Realtime-Whisper is OpenAI’s streaming speech-to-text model for live audio. Instead of waiting for a complete recording and returning one finished transcript, it can provide incremental transcript updates as speech arrives. This makes it suitable for applications where latency matters, such as live captions, meeting notes, call monitoring and voice-agent input.

The model is part of OpenAI’s realtime model lineup, but it has a narrower role than a speech-to-speech conversational model. GPT-Realtime-Whisper focuses on recognizing spoken language and returning text. It does not generate a spoken answer or act as a general-purpose assistant by itself.

OpenAI lists the model as available through its Realtime transcription infrastructure. The supplied release information gives May 7, 2026 as its launch date, while the official model documentation identifies it as a current model for realtime transcription.

How the realtime transcription workflow works

An application sends live audio to a transcription session and receives text updates as the speaker talks. A transcript delta is a partial update rather than a complete final response, so a user interface can append or revise displayed text as more speech is processed.

This design is different from uploading a finished audio file for an offline transcription job. A streaming workflow is useful when the text has value immediately: a caption may need to appear during a presentation, a meeting application may build notes continuously, or a voice agent may need recognized speech before deciding what to do next.

The model can also receive supporting text configuration. Its documented input and output profile is audio and text input with text output. Audio is the important input modality; the returned result is text that an application can display, store, search or pass to another system.

Key capabilities and supported modalities

  • Realtime speech-to-text: converts live audio into text during an active transcription session.
  • Streaming updates: provides low-latency transcript deltas while a person is speaking.
  • Text output: returns transcript content rather than spoken audio.
  • Audio input: accepts live audio as its principal content modality.
  • Realtime API access: is intended for OpenAI’s realtime transcription infrastructure.
  • Audio-duration pricing: is priced by minutes of audio rather than a conventional text-token input and output schedule.

GPT-Realtime-Whisper does not natively produce images, video, audio responses or embeddings. The supplied specifications also list function calling and structured outputs as unsupported. If an application needs to trigger tools, populate a strict JSON schema or perform business logic, those functions must be implemented around the transcription model by the surrounding application.

Pricing, context and output limits

The documented price is $0.017 per minute of audio. This is the primary pricing figure supplied for the model. The research describes the transcription price as covering the audio-duration workflow and does not document a separate text-output charge.

GPT-Realtime-Whisper has a listed 16,000-token context window and a 2,000-token maximum output. The context window represents the amount of text and conversational or session context the model can consider, while the output limit constrains the amount of text it can return in one output. These limits matter when a long-running application retains substantial transcript history or sends additional instructions alongside the audio.

A practical implementation should therefore manage session context rather than allowing an indefinitely growing transcript to accumulate. Applications may need to summarize, segment or otherwise manage older transcript material outside the model when sessions become lengthy. The supplied specifications do not define a maximum audio-session duration, so no specific duration limit should be assumed from the token figures alone.

Strengths and trade-offs

The model’s main strength is the combination of streaming behavior and a focused transcription role. A specialist transcription model can be a better fit than a general-purpose language model when the central requirement is to turn live speech into text quickly and continuously.

  • Responsiveness: transcript updates can appear while speech is still happening, which is important for captions and interactive voice workflows.
  • Clear operational purpose: the model is designed for transcription rather than open-ended conversation, reducing the need to use a broader model for a narrowly defined task.
  • Predictable billing unit: audio-duration pricing makes the main usage driver easy to relate to recording or streaming time.
  • Useful downstream format: text can be displayed, indexed, analyzed or passed to another model or application component.

Its trade-off is capability breadth. GPT-Realtime-Whisper is not intended to reason through complex questions, write software, call external functions or respond with speech. The supplied editorial ratings describe its speed as high and its cost position as moderately favorable, but those are editorial evaluations rather than provider-published benchmark results. The model’s low reasoning and coding evaluations reflect its specialization, not a claim that it is designed for those tasks.

Best use cases

GPT-Realtime-Whisper is a good match when the application needs text before the audio session has finished. Suitable examples include:

  • Live captions: displaying speech as text during meetings, classes, presentations and events.
  • Meeting transcription: building a running transcript that can later be edited into notes or action items.
  • Accessibility: providing a text representation of spoken content for people who cannot easily hear or follow audio.
  • Call-center analysis: making live or near-live conversation text available to monitoring and analytics systems.
  • Voice-agent input: converting a user’s speech into text that a separate agent or application can interpret.
  • Interview and field recording workflows: producing searchable transcript content while an interview or session is underway.

In a voice-agent architecture, GPT-Realtime-Whisper can serve as the recognition layer: it hears the user and returns text, while another component handles intent detection, tool execution, reasoning and response generation. That separation can be useful when transcription and conversation management have different performance or compliance requirements.

When to choose GPT-Realtime-Whisper

Choose GPT-Realtime-Whisper when the primary requirement is low-latency transcription from live audio and the application can handle the rest of the workflow itself. It is especially appropriate when partial transcript updates are more valuable than waiting for a polished final transcript after recording ends.

It is also a sensible choice when audio duration is a more useful budgeting measure than text-token volume. At $0.017 per minute, teams can estimate a basic transcription cost from expected audio time, although actual application costs may also include other models, storage, networking and processing components.

Choose a different type of option when the task requires a spoken reply, a general conversation, tool execution or structured action output in the same model response. A speech-to-speech realtime model is more appropriate when the system must listen and talk back directly. A general reasoning or coding model is more suitable when the central task is analysis, software generation or multi-step problem solving. An offline transcription workflow may be preferable when immediate updates are unnecessary and the priority is a post-recording process rather than live interaction.

Limitations and unsupported features

GPT-Realtime-Whisper should not be treated as a complete voice assistant. Its documented role ends at streaming transcription. It does not natively generate spoken audio, images or video, and it is not described as an embedding model.

Function calling and structured outputs are listed as unsupported. Consequently, an application cannot rely on the model itself to produce provider-enforced tool calls or schema-constrained JSON. Developers who need those features must validate transcript text and implement tool selection, structured parsing or business rules in surrounding code. For safety-critical or high-impact workflows, transcript content should also be reviewed and validated before it triggers an external action.

The 16,000-token context window and 2,000-token maximum output should be considered during session design. Long meetings, calls or events may require segmentation and external transcript management. The supplied research does not provide accuracy benchmarks, language coverage, speaker-diarization specifications or a guaranteed latency figure, so those characteristics should be tested for the intended audio conditions instead of inferred from the model name.

Bottom line

GPT-Realtime-Whisper is a focused OpenAI model for turning live audio into streaming text. Its value comes from timely transcript updates, a dedicated realtime transcription workflow and audio-duration pricing. It is well suited to captions, meetings, call analysis, accessibility and voice-agent input pipelines.

The model is less appropriate when transcription is only one small part of a larger assistant that must reason, call tools or speak back. In those cases, GPT-Realtime-Whisper can still be useful as the speech-recognition component, but additional models and application logic will be needed to complete the experience.


Answers to Frequently Asked Questions

How does GPT-Realtime-Whisper differ from a general-purpose or speech-to-speech model?
GPT-Realtime-Whisper is specialized for low-latency transcription and provides streaming text updates. It does not reason through complex tasks, generate spoken replies or act as a complete voice assistant. A speech-to-speech model is better for direct listen-and-talk interactions, while a general reasoning model is better for analysis, coding and multi-step problem solving.
Can GPT-Realtime-Whisper generate speech or call tools?
No. GPT-Realtime-Whisper returns text from live audio and does not natively generate spoken responses, call functions or produce structured outputs. Tool execution, business logic, validation and speech generation must be handled by surrounding application components.
What are GPT-Realtime-Whisper’s context and output limits?
GPT-Realtime-Whisper has a listed 16,000-token context window and a 2,000-token maximum output. Long-running applications may need to segment transcripts, summarize older content or manage session history outside the model.
What is GPT-Realtime-Whisper used for?
GPT-Realtime-Whisper is a streaming speech-to-text model for converting live audio into incremental text updates. It is suitable for live captions, meeting transcription, call monitoring, accessibility features, interview workflows and voice-agent input.
How much does GPT-Realtime-Whisper cost?
GPT-Realtime-Whisper costs $0.017 per minute of audio. This is the documented audio-duration price, and the supplied information does not specify a separate text-output charge.


Sources 5
Provider

About OpenAI