GPT-Live-Transcribe

GPT-Live-Transcribe

by OpenAI · Current and generally available for realtime transcription

OpenAI's GPT-Live-Transcribe converts live audio into incremental and finalized text with tunable latency, contextual prompts, keyword hints, and language hints. It is optimized for realtime transcription rather than speech generation, diarization, reasoning, or completed-file batch processing.

Text Reasoning Coding
GPT-Live-Transcribe is designed for applications that need speech converted into text while audio is still arriving. It is intended for live captions, call transcription, microphone-based interfaces, telephony streams, and other realtime workflows rather than delayed transcription of completed recordings.
Outputs

What GPT-Live-Transcribe can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family GPT-Live-Transcribe
Model type Other
Release date 2026-07-28
Status Current and generally available for realtime transcription
Knowledge cutoff notes

OpenAI's public model documentation does not specify a knowledge cutoff for this speech-recognition model. Its behavior is based on incoming audio and supplied transcription context rather than a user-facing factual knowledge cutoff.

Model notes

GPT-Live-Transcribe is a specialized streaming ASR model. It returns transcript deltas as speech arrives and a final transcript after each committed audio turn. It supports tunable delay, free-form context, keyword hints, and multiple expected language hints. OpenAI states that it does not return word-level timestamps, speaker labels, or transcription confidence scores. In transcription sessions, server_vad and semantic_vad are not supported; applications should omit turn detection or set it to null and commit audio turns explicitly. The model is listed at $0.017 per minute of realtime audio. Editorial reasoning and coding scores are intentionally minimal because those capabilities are outside the model's purpose.

Cost

Model pricing

Input $0.017 per minute of realtime audio
Output Included in the per-minute realtime audio price; OpenAI does not list a separate output-token price
Model guide

GPT-Live-Transcribe: Realtime Speech-to-Text, Pricing and API Capabilities

GPT-Live-Transcribe is OpenAI's low-latency streaming speech-to-text model for live audio. It produces incremental transcript updates and finalized text from realtime audio sessions, with controls for delay, contextual prompts, keyword hints, and expected languages.

What is GPT-Live-Transcribe?

GPT-Live-Transcribe is OpenAI's specialized realtime speech-to-text model. It accepts a live audio stream and returns transcript updates as speech is received, followed by a finalized transcript when an audio turn is committed. The model is listed as generally available for realtime transcription and is documented for use through OpenAI's realtime transcription workflow.

Unlike a conversational speech model, GPT-Live-Transcribe does not respond with spoken audio or carry out a dialogue. Its job is narrower and more operationally focused: turn incoming speech into text quickly enough for an application to display, store, search, or process it while the conversation is still happening.

That makes it a model for live microphones, phone calls, media streams, captions, interviews, meetings, and voice-interface input. For completed recordings that can be uploaded and processed later, a batch or file-transcription model may be a better fit.

Where GPT-Live-Transcribe fits in OpenAI's lineup

GPT-Live-Transcribe belongs to OpenAI's speech and realtime model group rather than its general-purpose text reasoning models. Its primary output is text generated from audio, and its design prioritizes low latency and continuous delivery of partial results.

OpenAI distinguishes this model from GPT-Transcribe, which is intended for completed audio files and asynchronous transcription workflows. The practical distinction is timing: GPT-Live-Transcribe is useful when an application needs text during an active audio session, while a file-transcription model is more suitable when the recording already exists and immediate updates are not required.

The model was released on July 28, 2026, according to the supplied model information. Its documented purpose remains specialized: it is not presented as a general reasoning, coding, tool-use, or speech-generation model.

Core capabilities

  • Streaming transcription: The model emits incremental transcript deltas while audio is being received.
  • Finalized turns: After an audio turn is committed, the API returns a final transcript event for that turn.
  • Latency controls: Applications can select delay settings that favor earlier partial text or allow more context before emitting results.
  • Context prompts: A free-form prompt can describe the recording or provide domain context.
  • Keyword hints: Applications can supply product names, acronyms, people, and specialized vocabulary that may otherwise be difficult to recognize.
  • Language hints: A session can include multiple expected language hints to help guide transcription.
  • Realtime connectivity: OpenAI documents WebSocket and WebRTC connectivity for its realtime transcription workflow.

These controls are particularly useful when the audio contains uncommon names, technical terminology, customer-account language, or industry-specific abbreviations. For example, a support application could provide the names of its products and internal systems as keyword hints before streaming a call.

How the realtime transcription workflow works

An application creates a transcription session and selects gpt-live-transcribe as the transcription model. It then streams audio chunks into an input audio buffer. When the client determines that a speech turn is complete, it commits that audio. The service can return partial transcription events during processing and a final transcript event after the turn is committed.

A transcript delta is an incremental piece of text, not necessarily a permanent sentence fragment. Later events can revise or extend earlier text as additional audio and context become available. User interfaces should therefore avoid treating every partial result as immutable. A common implementation approach is to associate events with their item identifiers, update the visible partial text as new deltas arrive, and replace it with the finalized transcript when the turn is complete.

The supplied documentation notes that server-side voice activity detection and semantic voice activity detection are not supported in the transcription-session configuration. Applications may need to detect turn boundaries on the client and commit audio explicitly. This is an important implementation consideration for microphone and telephony applications because the model does not automatically provide every part of the conversation-control layer.

Latency versus context

GPT-Live-Transcribe provides a delay setting that controls the trade-off between early output and additional context. Minimal and low-delay modes prioritize getting partial words or phrases to the application quickly. Medium, high, and extra-high settings allow more audio context before producing results and may improve transcription quality in some conditions.

There is no universally correct setting. A live caption display may prefer the earliest possible text, even if partial text is revised later. A contact-center application may accept more delay if it produces cleaner transcripts for archival search or agent assistance. The best choice depends on the target response time, recording quality, accents, background noise, speaking style, and vocabulary.

Context controls can complement the delay setting. A prompt can explain the subject of the recording, keyword hints can identify terms that must be recognized accurately, and expected language hints can guide multilingual sessions. These inputs do not remove the need to review important transcripts, but they give the model information that is not available from raw audio alone.

Supported inputs and outputs

CapabilityGPT-Live-Transcribe
Audio inputYes, through realtime audio streaming
Text inputContext can be supplied through prompts, keyword hints, language hints, and earlier transcript context
Text outputYes, as incremental deltas and finalized transcript events
Spoken audio outputNo
Image or video inputNo
Tool or function callingNo
Structured output or JSON modeNo
StreamingYes

The model is therefore multimodal in the limited sense that it accepts audio and produces text, but it is not a general multimodal assistant. It does not analyze images or video, generate audio, or return a spoken response.

Limitations and missing features

GPT-Live-Transcribe is intentionally specialized. It does not provide general reasoning, coding assistance, web search, tool calling, or speech-to-speech conversation. Its reasoning and coding scores in the supplied model data are editorial indicators rather than provider-published benchmark results, and they should not be interpreted as evidence that the model is suitable for those tasks.

OpenAI's realtime transcription documentation also states that the model does not return word-level timestamps, speaker labels, or transcription confidence scores. This matters for applications that need to identify who said each sentence, synchronize words precisely with a recording, or show a numeric confidence value for each recognition decision.

Speaker separation, often called diarization, is especially important for meetings, interviews, and multi-person calls. Because speaker labels are not provided, an application requiring diarized transcripts should use a compatible alternative or add a separate speaker-identification workflow. Similarly, applications that need detailed time alignment should choose a service that explicitly supplies word-level timestamps.

The supplied information does not specify a context-window length or maximum output-token limit for GPT-Live-Transcribe. Those values should therefore be treated as unknown rather than assumed to match another OpenAI model.

GPT-Live-Transcribe pricing

OpenAI lists GPT-Live-Transcribe at $0.017 per minute of realtime audio. The supplied pricing information describes this as the realtime audio price and does not list a separate output-token charge.

For cost planning, the relevant unit is the amount of audio sent through realtime transcription. A team should estimate the number of audio minutes generated by its microphones, calls, or media streams and account for the possibility that sessions remain open longer than expected. The model's per-minute pricing can be worthwhile when immediate transcript updates are a product requirement, but a batch transcription option may be more economical for recordings that do not need live results.

Best use cases

  • Live captions: Displaying speech as it happens in an application, event interface, or accessibility workflow.
  • Contact-center transcription: Producing a live text stream from customer-support or sales calls.
  • Telephony streams: Converting phone audio into text for monitoring, search, or downstream processing.
  • Microphone interfaces: Capturing speech input for an application that needs text before the user finishes a longer session.
  • Meeting and interview capture: Showing a running transcript while a conversation is underway, provided the lack of speaker labels is acceptable.
  • Voice-interface input: Supplying text to another application component when spoken audio output is not required from this model.

For production systems, developers should design for partial-result revisions, explicit turn commits, unreliable or noisy audio, and the absence of confidence scores. Important transcripts should also be reviewed before being used for legal, medical, financial, compliance, or other high-impact decisions.

When to choose GPT-Live-Transcribe

Choose GPT-Live-Transcribe when the central requirement is low-latency speech recognition from an active audio stream. It is a good match when users need to see text during a call, microphone session, meeting, or broadcast rather than waiting for an entire file to finish processing.

Its main trade-off is specialization. Compared with a general-purpose language model, GPT-Live-Transcribe offers the right interface for streaming speech but does not reason over the transcript, call tools, generate spoken replies, or produce structured application actions. Those tasks would need to be handled by other components in the application.

Compared with a completed-file transcription workflow, it prioritizes immediacy over the convenience of processing a finished recording in one batch. Choose a file-oriented alternative when the audio is already recorded, live updates are unnecessary, or the workflow depends on features such as diarization, detailed timestamps, or confidence scores that this model does not provide.

In short, GPT-Live-Transcribe is best selected for realtime audio-to-text delivery. It is not the appropriate choice merely because an application contains audio; the deciding question is whether transcript text must arrive while the audio session is still in progress.


Answers to Frequently Asked Questions

What are the main limitations of GPT-Live-Transcribe?
GPT-Live-Transcribe produces text only and does not generate spoken audio, analyze images or video, call tools, return structured JSON, or provide general reasoning and coding assistance. It also does not provide word-level timestamps, speaker labels, diarization, or transcription confidence scores.
When should I choose GPT-Live-Transcribe instead of a file-transcription model?
Choose GPT-Live-Transcribe when transcript text must arrive during an active microphone session, phone call, meeting, broadcast, or other live audio stream. A file-transcription or batch model is generally a better fit for completed recordings when realtime updates are unnecessary or when features such as speaker labels, detailed timestamps, or confidence scores are required.
How much does GPT-Live-Transcribe cost?
GPT-Live-Transcribe is listed at $0.017 per minute of realtime audio. The supplied pricing information does not specify a separate output-token charge, so cost estimates should be based primarily on the number of audio minutes sent through the realtime transcription service.
What is GPT-Live-Transcribe?
GPT-Live-Transcribe is OpenAI's realtime speech-to-text model. It accepts a live audio stream and returns incremental transcript updates while speech is received, followed by a finalized transcript when an audio turn is committed.
How does GPT-Live-Transcribe work in a realtime application?
An application creates a realtime transcription session, selects the gpt-live-transcribe model, and streams audio chunks into an input buffer. The service returns partial transcript deltas during processing and a final transcript event after the application commits the completed speech turn. Applications should expect partial results to be revised and may need to detect turn boundaries on the client.


Sources 4
Provider

About OpenAI