What is GPT-Live-Transcribe?
GPT-Live-Transcribe is OpenAI's specialized realtime speech-to-text model. It accepts a live audio stream and returns transcript updates as speech is received, followed by a finalized transcript when an audio turn is committed. The model is listed as generally available for realtime transcription and is documented for use through OpenAI's realtime transcription workflow.
Unlike a conversational speech model, GPT-Live-Transcribe does not respond with spoken audio or carry out a dialogue. Its job is narrower and more operationally focused: turn incoming speech into text quickly enough for an application to display, store, search, or process it while the conversation is still happening.
That makes it a model for live microphones, phone calls, media streams, captions, interviews, meetings, and voice-interface input. For completed recordings that can be uploaded and processed later, a batch or file-transcription model may be a better fit.
Where GPT-Live-Transcribe fits in OpenAI's lineup
GPT-Live-Transcribe belongs to OpenAI's speech and realtime model group rather than its general-purpose text reasoning models. Its primary output is text generated from audio, and its design prioritizes low latency and continuous delivery of partial results.
OpenAI distinguishes this model from GPT-Transcribe, which is intended for completed audio files and asynchronous transcription workflows. The practical distinction is timing: GPT-Live-Transcribe is useful when an application needs text during an active audio session, while a file-transcription model is more suitable when the recording already exists and immediate updates are not required.
The model was released on July 28, 2026, according to the supplied model information. Its documented purpose remains specialized: it is not presented as a general reasoning, coding, tool-use, or speech-generation model.
Core capabilities
- Streaming transcription: The model emits incremental transcript deltas while audio is being received.
- Finalized turns: After an audio turn is committed, the API returns a final transcript event for that turn.
- Latency controls: Applications can select delay settings that favor earlier partial text or allow more context before emitting results.
- Context prompts: A free-form prompt can describe the recording or provide domain context.
- Keyword hints: Applications can supply product names, acronyms, people, and specialized vocabulary that may otherwise be difficult to recognize.
- Language hints: A session can include multiple expected language hints to help guide transcription.
- Realtime connectivity: OpenAI documents WebSocket and WebRTC connectivity for its realtime transcription workflow.
These controls are particularly useful when the audio contains uncommon names, technical terminology, customer-account language, or industry-specific abbreviations. For example, a support application could provide the names of its products and internal systems as keyword hints before streaming a call.
How the realtime transcription workflow works
An application creates a transcription session and selects gpt-live-transcribe as the transcription model. It then streams audio chunks into an input audio buffer. When the client determines that a speech turn is complete, it commits that audio. The service can return partial transcription events during processing and a final transcript event after the turn is committed.
A transcript delta is an incremental piece of text, not necessarily a permanent sentence fragment. Later events can revise or extend earlier text as additional audio and context become available. User interfaces should therefore avoid treating every partial result as immutable. A common implementation approach is to associate events with their item identifiers, update the visible partial text as new deltas arrive, and replace it with the finalized transcript when the turn is complete.
The supplied documentation notes that server-side voice activity detection and semantic voice activity detection are not supported in the transcription-session configuration. Applications may need to detect turn boundaries on the client and commit audio explicitly. This is an important implementation consideration for microphone and telephony applications because the model does not automatically provide every part of the conversation-control layer.
Latency versus context
GPT-Live-Transcribe provides a delay setting that controls the trade-off between early output and additional context. Minimal and low-delay modes prioritize getting partial words or phrases to the application quickly. Medium, high, and extra-high settings allow more audio context before producing results and may improve transcription quality in some conditions.
There is no universally correct setting. A live caption display may prefer the earliest possible text, even if partial text is revised later. A contact-center application may accept more delay if it produces cleaner transcripts for archival search or agent assistance. The best choice depends on the target response time, recording quality, accents, background noise, speaking style, and vocabulary.
Context controls can complement the delay setting. A prompt can explain the subject of the recording, keyword hints can identify terms that must be recognized accurately, and expected language hints can guide multilingual sessions. These inputs do not remove the need to review important transcripts, but they give the model information that is not available from raw audio alone.
Supported inputs and outputs
| Capability | GPT-Live-Transcribe |
|---|---|
| Audio input | Yes, through realtime audio streaming |
| Text input | Context can be supplied through prompts, keyword hints, language hints, and earlier transcript context |
| Text output | Yes, as incremental deltas and finalized transcript events |
| Spoken audio output | No |
| Image or video input | No |
| Tool or function calling | No |
| Structured output or JSON mode | No |
| Streaming | Yes |
The model is therefore multimodal in the limited sense that it accepts audio and produces text, but it is not a general multimodal assistant. It does not analyze images or video, generate audio, or return a spoken response.
Limitations and missing features
GPT-Live-Transcribe is intentionally specialized. It does not provide general reasoning, coding assistance, web search, tool calling, or speech-to-speech conversation. Its reasoning and coding scores in the supplied model data are editorial indicators rather than provider-published benchmark results, and they should not be interpreted as evidence that the model is suitable for those tasks.
OpenAI's realtime transcription documentation also states that the model does not return word-level timestamps, speaker labels, or transcription confidence scores. This matters for applications that need to identify who said each sentence, synchronize words precisely with a recording, or show a numeric confidence value for each recognition decision.
Speaker separation, often called diarization, is especially important for meetings, interviews, and multi-person calls. Because speaker labels are not provided, an application requiring diarized transcripts should use a compatible alternative or add a separate speaker-identification workflow. Similarly, applications that need detailed time alignment should choose a service that explicitly supplies word-level timestamps.
The supplied information does not specify a context-window length or maximum output-token limit for GPT-Live-Transcribe. Those values should therefore be treated as unknown rather than assumed to match another OpenAI model.
GPT-Live-Transcribe pricing
OpenAI lists GPT-Live-Transcribe at $0.017 per minute of realtime audio. The supplied pricing information describes this as the realtime audio price and does not list a separate output-token charge.
For cost planning, the relevant unit is the amount of audio sent through realtime transcription. A team should estimate the number of audio minutes generated by its microphones, calls, or media streams and account for the possibility that sessions remain open longer than expected. The model's per-minute pricing can be worthwhile when immediate transcript updates are a product requirement, but a batch transcription option may be more economical for recordings that do not need live results.
Best use cases
- Live captions: Displaying speech as it happens in an application, event interface, or accessibility workflow.
- Contact-center transcription: Producing a live text stream from customer-support or sales calls.
- Telephony streams: Converting phone audio into text for monitoring, search, or downstream processing.
- Microphone interfaces: Capturing speech input for an application that needs text before the user finishes a longer session.
- Meeting and interview capture: Showing a running transcript while a conversation is underway, provided the lack of speaker labels is acceptable.
- Voice-interface input: Supplying text to another application component when spoken audio output is not required from this model.
For production systems, developers should design for partial-result revisions, explicit turn commits, unreliable or noisy audio, and the absence of confidence scores. Important transcripts should also be reviewed before being used for legal, medical, financial, compliance, or other high-impact decisions.
When to choose GPT-Live-Transcribe
Choose GPT-Live-Transcribe when the central requirement is low-latency speech recognition from an active audio stream. It is a good match when users need to see text during a call, microphone session, meeting, or broadcast rather than waiting for an entire file to finish processing.
Its main trade-off is specialization. Compared with a general-purpose language model, GPT-Live-Transcribe offers the right interface for streaming speech but does not reason over the transcript, call tools, generate spoken replies, or produce structured application actions. Those tasks would need to be handled by other components in the application.
Compared with a completed-file transcription workflow, it prioritizes immediacy over the convenience of processing a finished recording in one batch. Choose a file-oriented alternative when the audio is already recorded, live updates are unnecessary, or the workflow depends on features such as diarization, detailed timestamps, or confidence scores that this model does not provide.
In short, GPT-Live-Transcribe is best selected for realtime audio-to-text delivery. It is not the appropriate choice merely because an application contains audio; the deciding question is whether transcript text must arrive while the audio session is still in progress.

