GPT-Transcribe

GPT-Transcribe

by OpenAI · Current

OpenAI's GPT-Transcribe is a focused speech-to-text model for recorded audio, streamed file transcripts, and committed Realtime turns. It supports contextual prompts, keyword hints, multiple language hints, detected-language information, and common audio or video containers up to 25 MB in the standard file workflow. The model costs $0.0045 per audio minute and does not provide general reasoning, coding, function calling, structured outputs, or audio generation.

Text Reasoning Coding
GPT-Transcribe is OpenAI's recommended model for accurate transcription of recorded speech in its original language. It accepts audio and returns text, making it suitable for meetings, interviews, support recordings, media workflows, multilingual audio, and speech containing specialized terminology. The model is priced by audio duration at $0.0045 per minute.
Outputs

What GPT-Transcribe can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Batch API
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-Transcribe
Model type Other
Release date 2026-07-28
Status Current
Knowledge cutoff notes

OpenAI's current GPT-Transcribe documentation does not publish a knowledge cutoff. As a transcription model, its principal behavior is based on processing supplied audio rather than answering general knowledge questions.

Model notes

Canonical model ID is gpt-transcribe. OpenAI describes it as the default high-accuracy speech-to-text model for completed audio files, streamed file transcripts, and committed turns in Realtime sessions over WebSocket. It supports unstructured context, keyword hints, and multiple language hints. Standard file uploads can be up to 25 MB, with support for mp3, mp4, mpeg, mpga, m4a, wav, and webm formats. The transcription guide distinguishes GPT-Transcribe from specialized workflows for speaker labels, word timestamps, subtitle formats, and translation into English. Public documentation does not specify a conventional context window, maximum output-token limit, or knowledge cutoff. Editorial scores are comparative estimates for a specialized transcription model, not provider-published benchmarks.

Cost

Model pricing

Input $0.0045 per audio minute
Output No separate output-token price; included in the per-minute transcription price
Model guide

GPT-Transcribe: OpenAI Speech-to-Text Model, Pricing and API Capabilities

GPT-Transcribe is OpenAI's current high-accuracy speech-to-text model for converting recorded or committed audio into text. It supports file transcription, streamed transcripts for completed recordings, and committed turns in Realtime transcription sessions, with optional contextual prompts, keyword hints, and language hints.

What is GPT-Transcribe?

GPT-Transcribe is an OpenAI speech-to-text model designed to convert supplied audio into written text. It is intended primarily for completed recordings rather than general conversation, reasoning, or content generation. OpenAI positions it as the default high-accuracy model for transcribing recorded speech in its original language.

The model can be used with the standard transcription workflow for uploaded files, with streaming transcription for completed audio, and with committed audio turns in Realtime transcription sessions over WebSocket. In practical terms, this means an application can submit a recording and receive a transcript, or receive transcript text progressively while a bounded recording is being processed.

GPT-Transcribe is not a general-purpose GPT model. Its output is text transcription rather than an answer generated from general world knowledge. This distinction matters when selecting it: the model is optimized for recognizing speech, including terminology supplied through context, rather than for chat, coding, autonomous tool use, or spoken responses.

Where GPT-Transcribe fits in OpenAI's lineup

GPT-Transcribe is part of OpenAI's audio and transcription offerings. It is separate from OpenAI's general conversational models and from models designed to generate speech, images, video, or embeddings. The supplied documentation describes it as the current high-accuracy transcription model and lists it as available through the API.

It also occupies a different position from live conversational audio systems. GPT-Transcribe can support committed turns in a Realtime transcription workflow, but developers should not automatically treat it as a complete live voice assistant. A live microphone or telephone application may require a Realtime workflow that handles ongoing audio capture, turn detection, and related session behavior.

OpenAI's documentation also distinguishes GPT-Transcribe from older GPT-4o transcription models, which are scheduled for removal from the API on February 26, 2027, according to the supplied research. GPT-Transcribe is therefore the more current choice when the requirement is high-accuracy transcription through OpenAI's available transcription API.

Inputs, outputs and supported formats

The model accepts audio input and produces text output. It can process common audio and video container formats through the standard transcription endpoint, including MP3, MP4, MPEG, MPGA, M4A, WAV, and WebM. The standard file-transcription workflow has a 25 MB upload limit.

Although some supported containers can hold video, GPT-Transcribe's role is audio transcription. It does not produce video or interpret the task as video understanding. The relevant input is the spoken audio contained in the submitted file.

CapabilityGPT-Transcribe support
Audio inputYes
Text outputYes
Image input or outputNo
Video outputNo
Audio or speech outputNo
Streaming transcriptionYes, for supported transcription workflows
Standard file upload limit25 MB

Applications can provide a prompt containing unstructured context, literal keyword hints, and one or more expected language hints. For this model, the current API uses the plural languages field rather than the older singular language parameter. Language hints can be useful for multilingual recordings or audio that switches between languages, while keyword hints can improve recognition of product names, account identifiers, technical terms, and proper nouns.

Accuracy, context and recognition control

The principal reason to choose GPT-Transcribe is recognition quality for recorded speech. OpenAI describes it as a high-accuracy model, but the supplied research does not provide a standardized benchmark score. Accuracy will also depend on the recording, background noise, microphone quality, accents, overlapping speakers, and the terminology used in the audio.

Contextual instructions give the application a way to explain what it expects before transcription begins. For example, a support platform could provide a list of product names or technical vocabulary likely to occur in a customer call. This does not turn the model into a domain expert; it gives the transcription process additional clues about terms that might otherwise be difficult to distinguish.

GPT-Transcribe can return detected-language information alongside the transcript. That information can help an application decide how to label or route a recording, although the model should still be evaluated with the languages and audio conditions used by the specific application.

OpenAI does not publish a conventional context-window size, maximum output-token limit, or knowledge cutoff for GPT-Transcribe. Those omissions are less unusual for a transcription service because pricing and processing are based on audio duration rather than ordinary text input and output tokens. Developers should nevertheless test long recordings, segmentation, response handling, and storage behavior in their own integration instead of assuming that a recording of any length can be submitted as one request.

Pricing and API access

GPT-Transcribe is priced at $0.0045 per minute of transcription audio. The supplied pricing information does not list a separate output-token charge; the transcription price covers the model's text result on a per-audio-minute basis.

This pricing structure makes the main cost driver easy to estimate: audio duration. A ten-minute recording would have a model transcription charge of approximately $0.045, before any other application, storage, networking, or service costs. Actual billing should be checked against OpenAI's current pricing documentation before deployment.

The model is available through OpenAI's transcription API and is documented for file transcription and Realtime transcription workflows. Streaming responses can emit transcript text deltas followed by a final completed transcript event. Applications that display live progress should distinguish these interim or incremental events from the final transcript used for storage or downstream processing.

GPT-Transcribe does not support function calling or structured outputs as model features. It is therefore not a suitable choice when the primary requirement is for the model itself to return validated JSON objects, invoke business tools, or execute actions. A separate application layer can process the returned text, but that is different from native structured-output or function-calling support.

Speed, cost and capability trade-offs

Editorial evaluation in the supplied model data rates GPT-Transcribe highly for speed and cost relative to the other models in the catalog, while giving it low reasoning and coding scores. These are comparative editorial estimates, not OpenAI-published benchmark results. They should be interpreted as a positioning aid rather than a guarantee of latency, accuracy, or price in every workload.

The practical trade-off is straightforward. GPT-Transcribe concentrates its capability on turning audio into text, avoiding the broader reasoning and tool-use features of general-purpose models. That specialization can make it a more appropriate and economical option for transcription than using a general conversational model for the same task. On the other hand, it cannot replace a reasoning model when the application must analyze the transcript, answer questions about it, write code, call tools, or make decisions.

A common architecture is to use GPT-Transcribe first and then pass the resulting text to another system for summarization, extraction, search, classification, or workflow automation. That two-stage design preserves the model's focused role and makes it easier to inspect the transcript before taking downstream actions.

Best use cases

GPT-Transcribe is a strong fit when the input is recorded or committed speech and the desired result is an accurate text transcript. Suitable examples include:

  • Meeting and interview transcription.
  • Customer-support and sales-call recordings.
  • Media and podcast transcription.
  • Multilingual recordings and audio that includes multiple expected languages.
  • Technical, medical, product, or account terminology supported by keyword hints.
  • Applications that need transcript text progressively for a completed recording.
  • Realtime workflows that need the final transcript of committed audio turns.

Keyword and context hints are particularly useful when a small recognition error could change meaning. Product names, customer identifiers, software libraries, organization names, and industry-specific vocabulary are all examples of terms an application may want to provide explicitly.

Limitations and when another option may be better

GPT-Transcribe should not be selected solely because an application handles audio. It produces text and does not generate spoken audio, images, video, embeddings, or direct actions. It is also not presented as a general-purpose conversational or reasoning model.

The model may not be the best option when speaker labeling, word-level timestamps, subtitle formats, or translation into English is the primary requirement. The supplied transcription guidance identifies specialized models or workflows for those needs. Before choosing GPT-Transcribe, confirm whether the application needs a plain transcript or a richer output format with speaker attribution, timing metadata, or translation.

It is also important to distinguish streaming a completed file from continuous live transcription. If the application captures an open-ended microphone stream or telephone conversation, use the documented Realtime transcription workflow rather than assuming that repeated file uploads provide the same behavior.

For questions that require reasoning over the transcript, use GPT-Transcribe as the speech-recognition stage and select a separate model or application component for analysis. Similarly, coding tasks should be handled by a coding-capable model after transcription, not by GPT-Transcribe itself. The supplied research gives GPT-Transcribe an editorial coding score of 1 and reasoning score of 2; these are comparative editorial assessments, not provider benchmarks, but they accurately reflect that coding and general reasoning are outside its primary purpose.

When to choose GPT-Transcribe

Choose GPT-Transcribe when you need OpenAI API transcription for recorded speech, want text rather than audio output, and value a focused high-accuracy workflow with optional language, context, and keyword hints. Its per-minute pricing is simple to estimate, and its support for streaming transcript events can help applications show progress without changing the final text-oriented purpose of the model.

Choose another option when the central requirement is live conversational interaction, speech generation, speaker diarization, precise word timestamps, subtitle production, English translation, structured JSON, tool calling, or post-transcription reasoning. GPT-Transcribe can be one component in those systems, but the supplied specifications do not support treating it as a complete solution for all of them.

Current status

As of September 23, 2026, GPT-Transcribe is listed as OpenAI's default high-accuracy transcription model and is accessible through the API. Its canonical model ID is gpt-transcribe. The documented release date is July 28, 2026. OpenAI has not published a conventional context limit, maximum output-token limit, or knowledge cutoff for the model, so those values should be treated as unavailable rather than inferred.


Answers to Frequently Asked Questions

What are the main limitations of GPT-Transcribe?
GPT-Transcribe produces text transcripts but does not generate speech, images, video, embeddings, validated JSON, or direct tool calls. It is also not intended for speaker diarization, word-level timestamps, subtitle creation, English translation, coding, or post-transcription reasoning without additional models or application components.
Does GPT-Transcribe support streaming transcription and real-time audio?
Yes. GPT-Transcribe supports streaming transcription for completed audio and committed audio turns in Realtime transcription sessions over WebSocket. However, continuous microphone or telephone conversations may require a broader Realtime workflow that handles ongoing audio capture, turn detection, and session behavior.
What audio formats and file sizes does GPT-Transcribe support?
GPT-Transcribe supports common audio and video containers, including MP3, MP4, MPEG, MPGA, M4A, WAV, and WebM. The standard file-transcription workflow has a 25 MB upload limit, and the model transcribes the audio contained in a video file rather than producing video understanding or output.
What is GPT-Transcribe used for?
GPT-Transcribe is an OpenAI speech-to-text model designed to convert recorded or committed audio into written text. It is suitable for meetings, interviews, customer-support calls, podcasts, media, and multilingual recordings, but it is not a general-purpose reasoning or conversational model.
How much does GPT-Transcribe cost?
GPT-Transcribe costs $0.0045 per minute of transcription audio. For example, transcribing a 10-minute recording costs approximately $0.045 in model usage, excluding storage, networking, and other application costs.


Sources 5
Provider

About OpenAI