What is GPT-4o Transcribe?
GPT-4o Transcribe is OpenAI’s speech-to-text model for converting spoken audio into written text. Speech-to-text, also called automatic speech recognition, turns an audio recording or audio stream into a transcript that software can search, display, summarize, or pass to another model.
OpenAI introduced the model on March 20, 2025, as part of its next-generation audio model launch. It belongs to the GPT-4o family, but it is not a general-purpose GPT model for chat, reasoning, coding, or image generation. Its purpose is narrower: recognize speech and produce a text transcription.
The model is designed for applications such as meeting notes, customer-support call records, interviews, lectures, voice-agent input, and other workflows where accurate recognition of spoken language matters. It is currently accessible through OpenAI’s API, although OpenAI has deprecated it and scheduled its removal for February 26, 2027.
How GPT-4o Transcribe handles audio and prompts
The primary input is audio. Applications can send a recorded file through the Audio Transcriptions API, or use supported streaming and realtime transcription workflows. The output is text containing the recognized speech. Supported response formats also include text and JSON representations, which can make the result easier for software to process.
GPT-4o Transcribe can also accept an optional text prompt. This prompt gives the model context about the recording or instructions about how the transcript should be formatted. For example, a developer might provide product names, technical vocabulary, acronyms, expected capitalization, punctuation preferences, or context from an earlier audio segment. Supplying the expected language can improve transcription accuracy and latency.
- Input: Audio, with optional text prompting
- Output: Text transcription
- File processing: Supported through the Audio API
- Streaming: Supported for completed audio recordings and supported realtime workflows
- Response formats: Text and JSON formats are supported
The model does not natively produce speech, music, images, or video. It also does not function as a general conversational assistant. Its multimodal capability is limited to accepting audio and returning text.
Accuracy, language recognition, and prompting
OpenAI says GPT-4o Transcribe provides lower word error rates and better language recognition than its original Whisper models. A word error rate measures transcription mistakes against a reference transcript; a lower rate generally means fewer substitutions, omissions, or incorrectly inserted words. OpenAI’s launch material also highlights improved multilingual performance and handling of difficult conditions such as accents, background noise, and different speaking speeds.
These are provider claims rather than a guarantee for every recording. Real-world accuracy can still vary with microphone quality, overlapping speakers, background noise, pronunciation, code-switching, and specialized vocabulary. Important transcripts should therefore be reviewed before they are used for legal, medical, financial, or other high-consequence purposes.
Prompting is one of the model’s practical advantages. A prompt can help preserve the spelling of a company or product name, identify industry terminology, request consistent punctuation, or provide continuity between segments. This is particularly useful when a recording contains names and terms that a general speech recognizer might otherwise misinterpret.
Context window, output limit, and pricing
GPT-4o Transcribe has a 16,000-token context window and a maximum output length of 2,000 tokens. The context window is the amount of model input and working context available for a request, while the output limit restricts how much transcript text the model can return in one response. Long recordings may therefore need to be divided into segments, especially when the transcript would exceed the maximum output.
OpenAI’s listed pricing is:
| Usage type | Price |
|---|---|
| Audio input | $2.50 per 1 million audio tokens |
| Audio output | $10.00 per 1 million audio tokens |
Audio-token billing is different from charging a simple per-minute recording fee. The final cost depends on how the audio is represented and how many input and output tokens the request uses. Developers should estimate costs using their expected recording volume and verify the current figures in OpenAI’s model documentation before deploying a large transcription workload.
API access and deprecation status
GPT-4o Transcribe is available through OpenAI’s transcription-related API surfaces, including the Audio Transcriptions API and supported realtime transcription sessions. Existing integrations can process uploaded recordings and receive streaming partial transcript output for supported completed-audio workflows.
However, the model is not a long-term choice for a new production system. OpenAI announced its deprecation on August 26, 2026, and scheduled API removal for February 26, 2027. The supplied model information identifies GPT-Live-Transcribe and GPT-Transcribe as recommended future alternatives. Developers already using GPT-4o Transcribe should plan migration, test transcript quality with the replacement model, and confirm that their chosen endpoint supports the same response and streaming behavior they need.
For a current alternative focused on transcription, see the separate GPT-Transcribe model article. That model should be evaluated as an alternative rather than treated as an interchangeable drop-in replacement without testing.
Important limitations
GPT-4o Transcribe does not provide native speaker diarization. Diarization means identifying which speaker is talking at each point in a recording and adding labels such as “Speaker 1” and “Speaker 2.” OpenAI offers a separate GPT-4o Transcribe Diarize model for that purpose. If a meeting or interview requires reliable speaker labels, the diarization-specific model is the more appropriate option.
The model also lacks the broader capabilities associated with general-purpose language models. It is not intended for open-ended chat, software development, image generation, speech synthesis, or complex reasoning after the transcript has been produced. Those tasks may require a separate model in a multi-step workflow: GPT-4o Transcribe can create the transcript, and another system can summarize, classify, extract information, or answer questions about it.
Its scheduled shutdown is the most significant practical limitation. Even if its accuracy and prompting behavior meet current requirements, building a new dependency on a deprecated model creates migration work and future availability risk.
Reasoning, coding, tools, and performance trade-offs
GPT-4o Transcribe has little direct reasoning or coding capability because transcription is its specialized task. It recognizes and formats speech; it does not independently solve a programming problem, browse the web, or call external functions as a general-purpose assistant. Its tool-use score is listed as zero in the supplied model data, and no web-search capability is specified.
The specialization can still be useful for performance and cost. A transcription model is a better fit than a general chat model when the required operation is simply converting audio to text. It avoids paying for capabilities that are not needed, while prompting can improve results for domain-specific vocabulary. The trade-off is that it cannot replace a broader model when the application needs transcript analysis, multi-step reasoning, coding, or conversational responses.
Editorially, the supplied evaluation rates its speed highly and its cost efficiency as moderate, but these are catalog scores rather than OpenAI-published benchmark facts. Actual latency and total cost depend on audio length, request size, streaming behavior, network conditions, and the amount of downstream processing.
When to choose GPT-4o Transcribe
GPT-4o Transcribe may be suitable when an existing OpenAI integration needs accurate audio-to-text conversion and the project can operate during the model’s remaining availability window. It is especially relevant for:
- Meeting, interview, lecture, and customer-support transcription
- Voice-agent input that must be converted into text before further processing
- Recordings containing technical terms, product names, or industry vocabulary that benefit from prompting
- Applications that need file transcription, streaming output, or supported realtime transcription sessions
- Teams maintaining an existing integration that have not yet completed migration
For a new application, choose it only after considering the shutdown date and comparing the recommended replacement models. GPT-Transcribe or GPT-Live-Transcribe may be more appropriate for a new long-term deployment, while a diarization-specific model is preferable when speaker labels are central to the workflow. A general-purpose model should be added after transcription when the application needs summarization, extraction, reasoning, or conversation rather than recognition alone.
Overall assessment
GPT-4o Transcribe is a focused OpenAI speech-recognition model with audio input, text output, prompt-based terminology guidance, file support, streaming workflows, a 16,000-token context window, and documented token-based pricing. OpenAI reports better recognition and multilingual performance than its original Whisper models, making it useful for accurate transcription workloads.
Its limitations are equally important: it does not generate audio or other media, does not provide native speaker diarization, has limited relevance for reasoning and coding, and is scheduled for API shutdown on February 26, 2027. It remains a reasonable model to understand and potentially use in an existing system, but new projects should normally evaluate OpenAI’s designated successor models before committing to it.

