What is GPT-Transcribe?
GPT-Transcribe is an OpenAI speech-to-text model designed to convert supplied audio into written text. It is intended primarily for completed recordings rather than general conversation, reasoning, or content generation. OpenAI positions it as the default high-accuracy model for transcribing recorded speech in its original language.
The model can be used with the standard transcription workflow for uploaded files, with streaming transcription for completed audio, and with committed audio turns in Realtime transcription sessions over WebSocket. In practical terms, this means an application can submit a recording and receive a transcript, or receive transcript text progressively while a bounded recording is being processed.
GPT-Transcribe is not a general-purpose GPT model. Its output is text transcription rather than an answer generated from general world knowledge. This distinction matters when selecting it: the model is optimized for recognizing speech, including terminology supplied through context, rather than for chat, coding, autonomous tool use, or spoken responses.
Where GPT-Transcribe fits in OpenAI's lineup
GPT-Transcribe is part of OpenAI's audio and transcription offerings. It is separate from OpenAI's general conversational models and from models designed to generate speech, images, video, or embeddings. The supplied documentation describes it as the current high-accuracy transcription model and lists it as available through the API.
It also occupies a different position from live conversational audio systems. GPT-Transcribe can support committed turns in a Realtime transcription workflow, but developers should not automatically treat it as a complete live voice assistant. A live microphone or telephone application may require a Realtime workflow that handles ongoing audio capture, turn detection, and related session behavior.
OpenAI's documentation also distinguishes GPT-Transcribe from older GPT-4o transcription models, which are scheduled for removal from the API on February 26, 2027, according to the supplied research. GPT-Transcribe is therefore the more current choice when the requirement is high-accuracy transcription through OpenAI's available transcription API.
Inputs, outputs and supported formats
The model accepts audio input and produces text output. It can process common audio and video container formats through the standard transcription endpoint, including MP3, MP4, MPEG, MPGA, M4A, WAV, and WebM. The standard file-transcription workflow has a 25 MB upload limit.
Although some supported containers can hold video, GPT-Transcribe's role is audio transcription. It does not produce video or interpret the task as video understanding. The relevant input is the spoken audio contained in the submitted file.
| Capability | GPT-Transcribe support |
|---|---|
| Audio input | Yes |
| Text output | Yes |
| Image input or output | No |
| Video output | No |
| Audio or speech output | No |
| Streaming transcription | Yes, for supported transcription workflows |
| Standard file upload limit | 25 MB |
Applications can provide a prompt containing unstructured context, literal keyword hints, and one or more expected language hints. For this model, the current API uses the plural languages field rather than the older singular language parameter. Language hints can be useful for multilingual recordings or audio that switches between languages, while keyword hints can improve recognition of product names, account identifiers, technical terms, and proper nouns.
Accuracy, context and recognition control
The principal reason to choose GPT-Transcribe is recognition quality for recorded speech. OpenAI describes it as a high-accuracy model, but the supplied research does not provide a standardized benchmark score. Accuracy will also depend on the recording, background noise, microphone quality, accents, overlapping speakers, and the terminology used in the audio.
Contextual instructions give the application a way to explain what it expects before transcription begins. For example, a support platform could provide a list of product names or technical vocabulary likely to occur in a customer call. This does not turn the model into a domain expert; it gives the transcription process additional clues about terms that might otherwise be difficult to distinguish.
GPT-Transcribe can return detected-language information alongside the transcript. That information can help an application decide how to label or route a recording, although the model should still be evaluated with the languages and audio conditions used by the specific application.
OpenAI does not publish a conventional context-window size, maximum output-token limit, or knowledge cutoff for GPT-Transcribe. Those omissions are less unusual for a transcription service because pricing and processing are based on audio duration rather than ordinary text input and output tokens. Developers should nevertheless test long recordings, segmentation, response handling, and storage behavior in their own integration instead of assuming that a recording of any length can be submitted as one request.
Pricing and API access
GPT-Transcribe is priced at $0.0045 per minute of transcription audio. The supplied pricing information does not list a separate output-token charge; the transcription price covers the model's text result on a per-audio-minute basis.
This pricing structure makes the main cost driver easy to estimate: audio duration. A ten-minute recording would have a model transcription charge of approximately $0.045, before any other application, storage, networking, or service costs. Actual billing should be checked against OpenAI's current pricing documentation before deployment.
The model is available through OpenAI's transcription API and is documented for file transcription and Realtime transcription workflows. Streaming responses can emit transcript text deltas followed by a final completed transcript event. Applications that display live progress should distinguish these interim or incremental events from the final transcript used for storage or downstream processing.
GPT-Transcribe does not support function calling or structured outputs as model features. It is therefore not a suitable choice when the primary requirement is for the model itself to return validated JSON objects, invoke business tools, or execute actions. A separate application layer can process the returned text, but that is different from native structured-output or function-calling support.
Speed, cost and capability trade-offs
Editorial evaluation in the supplied model data rates GPT-Transcribe highly for speed and cost relative to the other models in the catalog, while giving it low reasoning and coding scores. These are comparative editorial estimates, not OpenAI-published benchmark results. They should be interpreted as a positioning aid rather than a guarantee of latency, accuracy, or price in every workload.
The practical trade-off is straightforward. GPT-Transcribe concentrates its capability on turning audio into text, avoiding the broader reasoning and tool-use features of general-purpose models. That specialization can make it a more appropriate and economical option for transcription than using a general conversational model for the same task. On the other hand, it cannot replace a reasoning model when the application must analyze the transcript, answer questions about it, write code, call tools, or make decisions.
A common architecture is to use GPT-Transcribe first and then pass the resulting text to another system for summarization, extraction, search, classification, or workflow automation. That two-stage design preserves the model's focused role and makes it easier to inspect the transcript before taking downstream actions.
Best use cases
GPT-Transcribe is a strong fit when the input is recorded or committed speech and the desired result is an accurate text transcript. Suitable examples include:
- Meeting and interview transcription.
- Customer-support and sales-call recordings.
- Media and podcast transcription.
- Multilingual recordings and audio that includes multiple expected languages.
- Technical, medical, product, or account terminology supported by keyword hints.
- Applications that need transcript text progressively for a completed recording.
- Realtime workflows that need the final transcript of committed audio turns.
Keyword and context hints are particularly useful when a small recognition error could change meaning. Product names, customer identifiers, software libraries, organization names, and industry-specific vocabulary are all examples of terms an application may want to provide explicitly.
Limitations and when another option may be better
GPT-Transcribe should not be selected solely because an application handles audio. It produces text and does not generate spoken audio, images, video, embeddings, or direct actions. It is also not presented as a general-purpose conversational or reasoning model.
The model may not be the best option when speaker labeling, word-level timestamps, subtitle formats, or translation into English is the primary requirement. The supplied transcription guidance identifies specialized models or workflows for those needs. Before choosing GPT-Transcribe, confirm whether the application needs a plain transcript or a richer output format with speaker attribution, timing metadata, or translation.
It is also important to distinguish streaming a completed file from continuous live transcription. If the application captures an open-ended microphone stream or telephone conversation, use the documented Realtime transcription workflow rather than assuming that repeated file uploads provide the same behavior.
For questions that require reasoning over the transcript, use GPT-Transcribe as the speech-recognition stage and select a separate model or application component for analysis. Similarly, coding tasks should be handled by a coding-capable model after transcription, not by GPT-Transcribe itself. The supplied research gives GPT-Transcribe an editorial coding score of 1 and reasoning score of 2; these are comparative editorial assessments, not provider benchmarks, but they accurately reflect that coding and general reasoning are outside its primary purpose.
When to choose GPT-Transcribe
Choose GPT-Transcribe when you need OpenAI API transcription for recorded speech, want text rather than audio output, and value a focused high-accuracy workflow with optional language, context, and keyword hints. Its per-minute pricing is simple to estimate, and its support for streaming transcript events can help applications show progress without changing the final text-oriented purpose of the model.
Choose another option when the central requirement is live conversational interaction, speech generation, speaker diarization, precise word timestamps, subtitle production, English translation, structured JSON, tool calling, or post-transcription reasoning. GPT-Transcribe can be one component in those systems, but the supplied specifications do not support treating it as a complete solution for all of them.
Current status
As of September 23, 2026, GPT-Transcribe is listed as OpenAI's default high-accuracy transcription model and is accessible through the API. Its canonical model ID is gpt-transcribe. The documented release date is July 28, 2026. OpenAI has not published a conventional context limit, maximum output-token limit, or knowledge cutoff for the model, so those values should be treated as unavailable rather than inferred.

