GPT-4o Transcribe

GPT-4o Transcribe Diarize

by OpenAI · Deprecated; currently accessible and scheduled for removal from the API on February 26, 2027.

Specialized OpenAI transcription model with built-in speaker diarization, diarized JSON output, speaker timestamps, optional speaker-reference clips, audio-token pricing, and a scheduled API shutdown on February 26, 2027.

Text Reasoning Coding
GPT-4o Transcribe Diarize combines automatic speech recognition with built-in speaker diarization. It is designed for meetings, interviews, calls, podcasts, and other recordings where a plain transcript is insufficient because the speakers must be distinguished. The model is currently accessible through OpenAI's Transcription API but is deprecated and scheduled for removal on February 26, 2027.
Outputs

What GPT-4o Transcribe Diarize can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o Transcribe
Model type Other
Context window 16K tokens
Maximum output 2K tokens
Knowledge cutoff June 1, 2024
Release date October 2025
Status Deprecated; currently accessible and scheduled for removal from the API on February 26, 2027.
Deprecation date 2026-08-26
Shutdown date 2027-02-26
Knowledge cutoff notes

The official model page lists June 1, 2024 as the model knowledge cutoff. This is separate from the date of the model's release and does not change when current audio is supplied for transcription.

Model notes

This is a specialized automatic speech recognition model available through the Transcription API. Use response_format=diarized_json to receive speaker labels and segment timing. For recordings longer than 30 seconds, chunking_strategy must be set to auto or a voice-activity-detection configuration. Up to four known-speaker reference clips may be provided, with each clip between 2 and 10 seconds. Prompting is not supported. The model is currently accessible but was deprecated on August 26, 2026 and is scheduled for shutdown on February 26, 2027. OpenAI recommends GPT Live Transcribe or GPT Transcribe as migration targets.

Cost

Model pricing

Input $2.50 per 1M audio tokens
Output $10.00 per 1M audio tokens
Model guide

GPT-4o Transcribe Diarize: Speaker-Labeled Transcription, Pricing and API Status

GPT-4o Transcribe Diarize is OpenAI's specialized speech-recognition model for converting multi-speaker audio into text while identifying which speaker produced each transcript segment.

What is GPT-4o Transcribe Diarize?

GPT-4o Transcribe Diarize is an OpenAI automatic speech recognition model for turning recorded audio into text while separating the conversation by speaker. Speaker diarization means determining which voice produced each portion of a recording. Instead of returning one undifferentiated transcript, the model can return segments with speaker identifiers, start times, end times, and the corresponding text.

This makes the model particularly suitable for multi-person recordings. A meeting transcript, for example, can preserve who made each statement; an interview can distinguish the interviewer from the guest; and a customer-support recording can separate an agent's comments from a caller's responses.

The model is available through OpenAI's audio Transcription API with the model identifier gpt-4o-transcribe-diarize. It belongs to OpenAI's speech-to-text model lineup rather than its general-purpose text-generation models. For ordinary transcription without speaker separation, OpenAI also offers the related GPT-4o Transcribe.

How speaker diarization works

To receive speaker-separated results, an application uses the diarized_json response format. The response contains transcript segments associated with speaker labels and timing metadata. An application can use those fields to display a readable dialogue, build a meeting summary grouped by participant, search what a particular speaker said, or export a structured conversation record.

The labels identify detected speakers, but they do not necessarily provide their real names automatically. OpenAI supports optional reference clips for known speakers to help map detected voices to known names. Up to four short reference clips can be supplied, and each clip should be between two and ten seconds long. This is useful when an application already knows the likely participants, such as a scheduled interview with a host and guest.

Reference clips should be treated as an aid to speaker identification rather than an absolute guarantee. Recordings with background noise, overlapping speech, poor microphone quality, or substantial changes in a speaker's voice may still require review. The supplied research does not provide an accuracy benchmark, so production teams should validate transcripts against representative recordings before relying on them without human checking.

Inputs, outputs and response formats

GPT-4o Transcribe Diarize accepts audio and produces text. Supported audio file formats include FLAC, MP3, MP4, MPEG, MPGA, M4A, OGG, WAV, and WebM. Its output is not synthesized speech: the model does not generate audio, music, or voice responses.

The documented response choices include regular JSON, plain text, and diarized JSON. The diarized format is the important option for this model's main use case because it exposes speaker-segment information and timing. The model can also stream transcription results as a recording is processed. Streaming can provide transcript events as processing progresses, while diarized responses can include finalized speaker-segment events.

For live microphone or media-stream transcription, OpenAI recommends separate realtime transcription models rather than treating this file-transcription model as a complete live-conversation solution. A related alternative in the available catalog is GPT-Live-Transcribe, which is positioned for realtime speech-to-text use cases.

Workflow and file-length limits

For recordings longer than 30 seconds, OpenAI requires a chunking strategy. The documented options include auto or a voice-activity-detection configuration. Chunking divides a longer recording into manageable sections, while voice activity detection can help identify portions containing speech. Applications should plan how to preserve ordering and combine segment results when processing a long meeting or podcast.

The model's documented context window is 16,000 tokens, and its maximum output is 2,000 tokens. These are model specifications rather than a promise that every audio file can be transcribed in one request. Long recordings may need to be divided into chunks, and the resulting transcript may need to be assembled by the application.

Word-level timestamp configuration is not available for this model's diarized response path. The supported diarized output instead provides segment-level start and end times. That distinction matters for workflows such as subtitle alignment, search-term highlighting, or applications that require precise timing for every word.

Main capabilities and limitations

  • Audio input: The model processes recorded audio files in supported formats.
  • Speaker separation: It identifies and labels different speakers within a recording.
  • Segment timing: Diarized output includes segment start and end times.
  • Known-speaker references: Up to four reference clips can help associate detected voices with known participants.
  • Streaming: Transcription results can be streamed while a completed recording is processed.
  • No image or video understanding: The model does not accept image or video input.
  • No generated audio: It does not provide speech synthesis, music generation, or other audio output.
  • No general text generation: Its primary output is a transcript, not a general-purpose written response.
  • No prompting: Prompting is not supported for this diarization model.
  • No word-level timestamps: The diarized response path provides segment timing rather than configurable word-by-word timing.
  • No listed tool or function support: The supplied specifications identify no tool-use capability for this model.

These limitations make the model narrower than a general multimodal or language model. Its strength is the specific combination of transcription and speaker labeling, not broad reasoning, coding, web search, image analysis, or content generation.

Pricing and model specifications

SpecificationGPT-4o Transcribe Diarize
ProviderOpenAI
Model typeAutomatic speech recognition with speaker diarization
InputAudio
OutputText, including diarized JSON responses
Context length16,000 tokens
Maximum output2,000 tokens
Input price$2.50 per 1 million audio tokens
Output price$10.00 per 1 million audio tokens
StatusDeprecated; currently accessible according to the supplied research

The listed price is token-based rather than a simple per-minute rate. Actual usage depends on how audio is represented and processed, so teams should estimate costs from representative files and include any additional processing required to chunk and combine long recordings. The supplied research does not provide a separate free tier, minimum charge, or fixed per-hour conversion.

Speed, cost and capability trade-offs

The supplied editorial scoring rates the model's speed at 8 out of 10 and its cost at 7 out of 10. These are evaluation scores, not OpenAI-published benchmarks or guarantees. They indicate that the model is viewed as a relatively efficient choice for its specialized task, but they should not be interpreted as a measured latency or universal cost ranking.

Compared with a basic speech-to-text workflow, GPT-4o Transcribe Diarize adds speaker identification and segment timing, which can reduce the amount of custom post-processing needed for interviews and meetings. That additional structure is the reason to choose it over ordinary transcription. However, if an application only needs a single block of text, the diarization features may add unnecessary complexity.

Compared with realtime transcription options, this model is better suited to processing recorded files through the Transcription API. A realtime model may be more appropriate when text must appear continuously during a live conversation or media stream. Compared with a general-purpose language model, GPT-4o Transcribe Diarize is the more direct tool for producing a speaker-labeled transcript, but it is not the appropriate choice for broad reasoning, coding, web research, or post-transcription analysis by itself.

Best use cases

  • Meeting records: Preserve statements by participant before generating minutes or action items.
  • Interviews: Separate interviewer questions from guest answers and retain timing for review.
  • Customer-support calls: Distinguish agent and customer speech for quality review or searchable records.
  • Podcasts: Create speaker-labeled transcripts for editing, accessibility, or content repurposing.
  • Research interviews: Organize responses by participant while retaining segment timestamps.
  • Known-participant recordings: Use short reference clips when an application needs to map detected speakers to expected names.

For sensitive recordings, organizations should also consider their own retention, access-control, consent, and privacy requirements. The supplied model research describes technical behavior and pricing but does not establish a specific compliance certification or retention guarantee for this model.

When to choose this model

Choose GPT-4o Transcribe Diarize when the recording contains multiple speakers and the speaker identity of each segment is a core requirement. It is especially useful when an application would otherwise need to combine ordinary transcription with a separate diarization system. The model is also a reasonable fit when segment-level timestamps and optional known-speaker references are more valuable than word-level timing.

Choose a regular transcription model instead when speaker labels are unnecessary, when the workflow is optimized for simpler text output, or when the application needs capabilities outside speech recognition. Choose a realtime transcription option when audio arrives continuously and the user needs an immediate transcript. Choose a separate analysis or language model after transcription when the next task is summarization, classification, extraction, coding, or question answering.

New deployments should consider the model's lifecycle status before committing to it. The supplied research says that GPT-4o Transcribe Diarize was deprecated on August 26, 2026, remains accessible as of September 23, 2026, and is scheduled for API removal on February 26, 2027. OpenAI identifies GPT Live Transcribe or GPT Transcribe as migration targets. Therefore, this model may still be useful for an existing compatible workflow or short-term evaluation, but it is a poor long-term default for a new system unless the team has a clear migration plan.

Bottom line

GPT-4o Transcribe Diarize is a focused audio transcription model whose defining feature is speaker-labeled output. It supports common audio formats, diarized JSON, segment timestamps, streaming events, and up to four known-speaker reference clips. Its 16,000-token context window, 2,000-token maximum output, and token-based pricing are useful planning constraints, while the lack of prompting and word-level timestamps limits some advanced workflows.

Its strongest reason for selection is straightforward: it can turn a multi-speaker recording into a structured transcript without requiring speaker separation to be built as a completely separate step. Its most important reason for caution is equally clear: the model is deprecated and scheduled for removal, so new production systems should evaluate the named replacement options and design for migration.


Answers to Frequently Asked Questions

How much does GPT-4o Transcribe Diarize cost?
The listed price is $2.50 per 1 million audio input tokens and $10.00 per 1 million output tokens. Costs are token-based rather than a fixed per-minute rate, so actual usage depends on how audio is processed, including any chunking required for longer recordings.
Is GPT-4o Transcribe Diarize still available, and should new projects use it?
The model is listed as deprecated but remains accessible according to the supplied research. It was scheduled for API removal on February 26, 2027, with GPT Live Transcribe or GPT Transcribe identified as migration targets. New projects should evaluate those alternatives and have a clear migration plan before relying on GPT-4o Transcribe Diarize.
What audio formats and timestamp options does GPT-4o Transcribe Diarize support?
The model supports FLAC, MP3, MP4, MPEG, MPGA, M4A, OGG, WAV, and WebM audio files. Its diarized output provides segment-level start and end times, but configurable word-level timestamps are not available for the diarized response path.
What is GPT-4o Transcribe Diarize used for?
GPT-4o Transcribe Diarize converts recorded audio into text while separating the transcript by speaker. It is designed for meetings, interviews, customer-support calls, podcasts, and research recordings where speaker labels and segment timestamps are important.
How does speaker diarization work in GPT-4o Transcribe Diarize?
Use the diarized_json response format to receive transcript segments with speaker identifiers, start times, end times, and text. Up to four reference clips of two to ten seconds each can help associate detected voices with known participants, although noisy recordings, overlapping speech, and poor audio may still require human review.


Sources 5
Provider

About OpenAI