What is GPT-Realtime-Whisper?
GPT-Realtime-Whisper is OpenAI’s streaming speech-to-text model for live audio. Instead of waiting for a complete recording and returning one finished transcript, it can provide incremental transcript updates as speech arrives. This makes it suitable for applications where latency matters, such as live captions, meeting notes, call monitoring and voice-agent input.
The model is part of OpenAI’s realtime model lineup, but it has a narrower role than a speech-to-speech conversational model. GPT-Realtime-Whisper focuses on recognizing spoken language and returning text. It does not generate a spoken answer or act as a general-purpose assistant by itself.
OpenAI lists the model as available through its Realtime transcription infrastructure. The supplied release information gives May 7, 2026 as its launch date, while the official model documentation identifies it as a current model for realtime transcription.
How the realtime transcription workflow works
An application sends live audio to a transcription session and receives text updates as the speaker talks. A transcript delta is a partial update rather than a complete final response, so a user interface can append or revise displayed text as more speech is processed.
This design is different from uploading a finished audio file for an offline transcription job. A streaming workflow is useful when the text has value immediately: a caption may need to appear during a presentation, a meeting application may build notes continuously, or a voice agent may need recognized speech before deciding what to do next.
The model can also receive supporting text configuration. Its documented input and output profile is audio and text input with text output. Audio is the important input modality; the returned result is text that an application can display, store, search or pass to another system.
Key capabilities and supported modalities
- Realtime speech-to-text: converts live audio into text during an active transcription session.
- Streaming updates: provides low-latency transcript deltas while a person is speaking.
- Text output: returns transcript content rather than spoken audio.
- Audio input: accepts live audio as its principal content modality.
- Realtime API access: is intended for OpenAI’s realtime transcription infrastructure.
- Audio-duration pricing: is priced by minutes of audio rather than a conventional text-token input and output schedule.
GPT-Realtime-Whisper does not natively produce images, video, audio responses or embeddings. The supplied specifications also list function calling and structured outputs as unsupported. If an application needs to trigger tools, populate a strict JSON schema or perform business logic, those functions must be implemented around the transcription model by the surrounding application.
Pricing, context and output limits
The documented price is $0.017 per minute of audio. This is the primary pricing figure supplied for the model. The research describes the transcription price as covering the audio-duration workflow and does not document a separate text-output charge.
GPT-Realtime-Whisper has a listed 16,000-token context window and a 2,000-token maximum output. The context window represents the amount of text and conversational or session context the model can consider, while the output limit constrains the amount of text it can return in one output. These limits matter when a long-running application retains substantial transcript history or sends additional instructions alongside the audio.
A practical implementation should therefore manage session context rather than allowing an indefinitely growing transcript to accumulate. Applications may need to summarize, segment or otherwise manage older transcript material outside the model when sessions become lengthy. The supplied specifications do not define a maximum audio-session duration, so no specific duration limit should be assumed from the token figures alone.
Strengths and trade-offs
The model’s main strength is the combination of streaming behavior and a focused transcription role. A specialist transcription model can be a better fit than a general-purpose language model when the central requirement is to turn live speech into text quickly and continuously.
- Responsiveness: transcript updates can appear while speech is still happening, which is important for captions and interactive voice workflows.
- Clear operational purpose: the model is designed for transcription rather than open-ended conversation, reducing the need to use a broader model for a narrowly defined task.
- Predictable billing unit: audio-duration pricing makes the main usage driver easy to relate to recording or streaming time.
- Useful downstream format: text can be displayed, indexed, analyzed or passed to another model or application component.
Its trade-off is capability breadth. GPT-Realtime-Whisper is not intended to reason through complex questions, write software, call external functions or respond with speech. The supplied editorial ratings describe its speed as high and its cost position as moderately favorable, but those are editorial evaluations rather than provider-published benchmark results. The model’s low reasoning and coding evaluations reflect its specialization, not a claim that it is designed for those tasks.
Best use cases
GPT-Realtime-Whisper is a good match when the application needs text before the audio session has finished. Suitable examples include:
- Live captions: displaying speech as text during meetings, classes, presentations and events.
- Meeting transcription: building a running transcript that can later be edited into notes or action items.
- Accessibility: providing a text representation of spoken content for people who cannot easily hear or follow audio.
- Call-center analysis: making live or near-live conversation text available to monitoring and analytics systems.
- Voice-agent input: converting a user’s speech into text that a separate agent or application can interpret.
- Interview and field recording workflows: producing searchable transcript content while an interview or session is underway.
In a voice-agent architecture, GPT-Realtime-Whisper can serve as the recognition layer: it hears the user and returns text, while another component handles intent detection, tool execution, reasoning and response generation. That separation can be useful when transcription and conversation management have different performance or compliance requirements.
When to choose GPT-Realtime-Whisper
Choose GPT-Realtime-Whisper when the primary requirement is low-latency transcription from live audio and the application can handle the rest of the workflow itself. It is especially appropriate when partial transcript updates are more valuable than waiting for a polished final transcript after recording ends.
It is also a sensible choice when audio duration is a more useful budgeting measure than text-token volume. At $0.017 per minute, teams can estimate a basic transcription cost from expected audio time, although actual application costs may also include other models, storage, networking and processing components.
Choose a different type of option when the task requires a spoken reply, a general conversation, tool execution or structured action output in the same model response. A speech-to-speech realtime model is more appropriate when the system must listen and talk back directly. A general reasoning or coding model is more suitable when the central task is analysis, software generation or multi-step problem solving. An offline transcription workflow may be preferable when immediate updates are unnecessary and the priority is a post-recording process rather than live interaction.
Limitations and unsupported features
GPT-Realtime-Whisper should not be treated as a complete voice assistant. Its documented role ends at streaming transcription. It does not natively generate spoken audio, images or video, and it is not described as an embedding model.
Function calling and structured outputs are listed as unsupported. Consequently, an application cannot rely on the model itself to produce provider-enforced tool calls or schema-constrained JSON. Developers who need those features must validate transcript text and implement tool selection, structured parsing or business rules in surrounding code. For safety-critical or high-impact workflows, transcript content should also be reviewed and validated before it triggers an external action.
The 16,000-token context window and 2,000-token maximum output should be considered during session design. Long meetings, calls or events may require segmentation and external transcript management. The supplied research does not provide accuracy benchmarks, language coverage, speaker-diarization specifications or a guaranteed latency figure, so those characteristics should be tested for the intended audio conditions instead of inferred from the model name.
Bottom line
GPT-Realtime-Whisper is a focused OpenAI model for turning live audio into streaming text. Its value comes from timely transcript updates, a dedicated realtime transcription workflow and audio-duration pricing. It is well suited to captions, meetings, call analysis, accessibility and voice-agent input pipelines.
The model is less appropriate when transcription is only one small part of a larger assistant that must reason, call tools or speak back. In those cases, GPT-Realtime-Whisper can still be useful as the speech-recognition component, but additional models and application logic will be needed to complete the experience.

