GPT-4o

GPT-4o Transcribe

by OpenAI · Deprecated; currently accessible; scheduled for API shutdown on 2027-02-26

OpenAI’s GPT-4o Transcribe converts audio into text with improved word error rates and language recognition compared with Whisper. It supports prompted transcription, file processing, streaming workflows, and realtime transcription, but is deprecated and scheduled for API removal on February 26, 2027.

Text Reasoning Coding
GPT-4o Transcribe is an OpenAI audio model built for developers who need accurate speech recognition rather than general-purpose conversation or content generation. It can transcribe recordings, live or completed audio streams, meetings, interviews, calls, lectures, and voice-agent input. The model accepts audio plus optional text instructions and returns text transcripts. Its listed price is $2.50 per one million audio input tokens and $10 per one million audio output tokens, but its scheduled shutdown makes OpenAI’s recommended replacement models more appropriate for most new projects.
Outputs

What GPT-4o Transcribe can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming JSON mode
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o
Model type Other
Context window 16K tokens
Maximum output 2K tokens
Knowledge cutoff June 1, 2024
Release date 2025-03-20
Status Deprecated; currently accessible; scheduled for API shutdown on 2027-02-26
Deprecation date 2026-08-26
Shutdown date 2027-02-26
Knowledge cutoff notes

The official model page lists June 1, 2024 as the model’s knowledge cutoff. This cutoff describes the underlying model and is not changed by transcription prompts, realtime sessions, or external application context.

Model notes

GPT-4o Transcribe is a specialized speech-recognition model, not a general-purpose GPT model. It accepts audio and optional text prompts and returns text transcripts. OpenAI reports improved word error rates and language recognition compared with original Whisper models. Existing integrations support file transcription, streaming transcription for completed recordings, and realtime transcription workflows. Speaker diarization is provided by the separate GPT-4o Transcribe Diarize model. OpenAI announced deprecation on August 26, 2026, and scheduled removal from the API for February 26, 2027, recommending GPT-Live-Transcribe or GPT-Transcribe as replacements.

Cost

Model pricing

Input $2.50 per 1M audio input tokens
Output $10.00 per 1M audio output tokens
Model guide

GPT-4o Transcribe: Pricing, Features, Limits, and API Status

GPT-4o Transcribe is OpenAI’s specialized speech-to-text model for converting recorded or streamed audio into text. It accepts audio and optional text prompts, supports file and streaming transcription workflows, and offers improved word error rates and language recognition compared with OpenAI’s original Whisper models. The model remains accessible but is deprecated and scheduled for API shutdown on February 26, 2027.

What is GPT-4o Transcribe?

GPT-4o Transcribe is OpenAI’s speech-to-text model for converting spoken audio into written text. Speech-to-text, also called automatic speech recognition, turns an audio recording or audio stream into a transcript that software can search, display, summarize, or pass to another model.

OpenAI introduced the model on March 20, 2025, as part of its next-generation audio model launch. It belongs to the GPT-4o family, but it is not a general-purpose GPT model for chat, reasoning, coding, or image generation. Its purpose is narrower: recognize speech and produce a text transcription.

The model is designed for applications such as meeting notes, customer-support call records, interviews, lectures, voice-agent input, and other workflows where accurate recognition of spoken language matters. It is currently accessible through OpenAI’s API, although OpenAI has deprecated it and scheduled its removal for February 26, 2027.

How GPT-4o Transcribe handles audio and prompts

The primary input is audio. Applications can send a recorded file through the Audio Transcriptions API, or use supported streaming and realtime transcription workflows. The output is text containing the recognized speech. Supported response formats also include text and JSON representations, which can make the result easier for software to process.

GPT-4o Transcribe can also accept an optional text prompt. This prompt gives the model context about the recording or instructions about how the transcript should be formatted. For example, a developer might provide product names, technical vocabulary, acronyms, expected capitalization, punctuation preferences, or context from an earlier audio segment. Supplying the expected language can improve transcription accuracy and latency.

  • Input: Audio, with optional text prompting
  • Output: Text transcription
  • File processing: Supported through the Audio API
  • Streaming: Supported for completed audio recordings and supported realtime workflows
  • Response formats: Text and JSON formats are supported

The model does not natively produce speech, music, images, or video. It also does not function as a general conversational assistant. Its multimodal capability is limited to accepting audio and returning text.

Accuracy, language recognition, and prompting

OpenAI says GPT-4o Transcribe provides lower word error rates and better language recognition than its original Whisper models. A word error rate measures transcription mistakes against a reference transcript; a lower rate generally means fewer substitutions, omissions, or incorrectly inserted words. OpenAI’s launch material also highlights improved multilingual performance and handling of difficult conditions such as accents, background noise, and different speaking speeds.

These are provider claims rather than a guarantee for every recording. Real-world accuracy can still vary with microphone quality, overlapping speakers, background noise, pronunciation, code-switching, and specialized vocabulary. Important transcripts should therefore be reviewed before they are used for legal, medical, financial, or other high-consequence purposes.

Prompting is one of the model’s practical advantages. A prompt can help preserve the spelling of a company or product name, identify industry terminology, request consistent punctuation, or provide continuity between segments. This is particularly useful when a recording contains names and terms that a general speech recognizer might otherwise misinterpret.

Context window, output limit, and pricing

GPT-4o Transcribe has a 16,000-token context window and a maximum output length of 2,000 tokens. The context window is the amount of model input and working context available for a request, while the output limit restricts how much transcript text the model can return in one response. Long recordings may therefore need to be divided into segments, especially when the transcript would exceed the maximum output.

OpenAI’s listed pricing is:

Usage typePrice
Audio input$2.50 per 1 million audio tokens
Audio output$10.00 per 1 million audio tokens

Audio-token billing is different from charging a simple per-minute recording fee. The final cost depends on how the audio is represented and how many input and output tokens the request uses. Developers should estimate costs using their expected recording volume and verify the current figures in OpenAI’s model documentation before deploying a large transcription workload.

API access and deprecation status

GPT-4o Transcribe is available through OpenAI’s transcription-related API surfaces, including the Audio Transcriptions API and supported realtime transcription sessions. Existing integrations can process uploaded recordings and receive streaming partial transcript output for supported completed-audio workflows.

However, the model is not a long-term choice for a new production system. OpenAI announced its deprecation on August 26, 2026, and scheduled API removal for February 26, 2027. The supplied model information identifies GPT-Live-Transcribe and GPT-Transcribe as recommended future alternatives. Developers already using GPT-4o Transcribe should plan migration, test transcript quality with the replacement model, and confirm that their chosen endpoint supports the same response and streaming behavior they need.

For a current alternative focused on transcription, see the separate GPT-Transcribe model article. That model should be evaluated as an alternative rather than treated as an interchangeable drop-in replacement without testing.

Important limitations

GPT-4o Transcribe does not provide native speaker diarization. Diarization means identifying which speaker is talking at each point in a recording and adding labels such as “Speaker 1” and “Speaker 2.” OpenAI offers a separate GPT-4o Transcribe Diarize model for that purpose. If a meeting or interview requires reliable speaker labels, the diarization-specific model is the more appropriate option.

The model also lacks the broader capabilities associated with general-purpose language models. It is not intended for open-ended chat, software development, image generation, speech synthesis, or complex reasoning after the transcript has been produced. Those tasks may require a separate model in a multi-step workflow: GPT-4o Transcribe can create the transcript, and another system can summarize, classify, extract information, or answer questions about it.

Its scheduled shutdown is the most significant practical limitation. Even if its accuracy and prompting behavior meet current requirements, building a new dependency on a deprecated model creates migration work and future availability risk.

Reasoning, coding, tools, and performance trade-offs

GPT-4o Transcribe has little direct reasoning or coding capability because transcription is its specialized task. It recognizes and formats speech; it does not independently solve a programming problem, browse the web, or call external functions as a general-purpose assistant. Its tool-use score is listed as zero in the supplied model data, and no web-search capability is specified.

The specialization can still be useful for performance and cost. A transcription model is a better fit than a general chat model when the required operation is simply converting audio to text. It avoids paying for capabilities that are not needed, while prompting can improve results for domain-specific vocabulary. The trade-off is that it cannot replace a broader model when the application needs transcript analysis, multi-step reasoning, coding, or conversational responses.

Editorially, the supplied evaluation rates its speed highly and its cost efficiency as moderate, but these are catalog scores rather than OpenAI-published benchmark facts. Actual latency and total cost depend on audio length, request size, streaming behavior, network conditions, and the amount of downstream processing.

When to choose GPT-4o Transcribe

GPT-4o Transcribe may be suitable when an existing OpenAI integration needs accurate audio-to-text conversion and the project can operate during the model’s remaining availability window. It is especially relevant for:

  • Meeting, interview, lecture, and customer-support transcription
  • Voice-agent input that must be converted into text before further processing
  • Recordings containing technical terms, product names, or industry vocabulary that benefit from prompting
  • Applications that need file transcription, streaming output, or supported realtime transcription sessions
  • Teams maintaining an existing integration that have not yet completed migration

For a new application, choose it only after considering the shutdown date and comparing the recommended replacement models. GPT-Transcribe or GPT-Live-Transcribe may be more appropriate for a new long-term deployment, while a diarization-specific model is preferable when speaker labels are central to the workflow. A general-purpose model should be added after transcription when the application needs summarization, extraction, reasoning, or conversation rather than recognition alone.

Overall assessment

GPT-4o Transcribe is a focused OpenAI speech-recognition model with audio input, text output, prompt-based terminology guidance, file support, streaming workflows, a 16,000-token context window, and documented token-based pricing. OpenAI reports better recognition and multilingual performance than its original Whisper models, making it useful for accurate transcription workloads.

Its limitations are equally important: it does not generate audio or other media, does not provide native speaker diarization, has limited relevance for reasoning and coding, and is scheduled for API shutdown on February 26, 2027. It remains a reasonable model to understand and potentially use in an existing system, but new projects should normally evaluate OpenAI’s designated successor models before committing to it.


Answers to Frequently Asked Questions

Does GPT-4o Transcribe support speaker diarization?
No. GPT-4o Transcribe does not natively identify and label individual speakers. Applications that need speaker labels should use OpenAI’s GPT-4o Transcribe Diarize model or another diarization-specific solution.
Is GPT-4o Transcribe still available through the API?
GPT-4o Transcribe is currently available through OpenAI’s transcription and supported realtime API workflows, but it has been deprecated. OpenAI has scheduled its API removal for February 26, 2027, so existing users should plan migration and new projects should evaluate replacement models.
What are GPT-4o Transcribe’s context and output limits?
GPT-4o Transcribe has a 16,000-token context window and a maximum output length of 2,000 tokens. Long recordings may need to be split into smaller segments, particularly when the transcript exceeds the output limit.
How much does GPT-4o Transcribe cost?
OpenAI lists pricing at $2.50 per 1 million audio input tokens and $10.00 per 1 million audio output tokens. Costs are based on audio-token usage rather than a simple per-minute fee, so the final price depends on how the audio is represented and processed.
What is GPT-4o Transcribe used for?
GPT-4o Transcribe is an OpenAI speech-to-text model that converts spoken audio into written text. It is designed for meeting notes, interviews, lectures, customer-support calls, voice-agent input, and other transcription workflows.


Sources 5
Provider

About OpenAI