What is GPT-4o Transcribe Diarize?
GPT-4o Transcribe Diarize is an OpenAI automatic speech recognition model for turning recorded audio into text while separating the conversation by speaker. Speaker diarization means determining which voice produced each portion of a recording. Instead of returning one undifferentiated transcript, the model can return segments with speaker identifiers, start times, end times, and the corresponding text.
This makes the model particularly suitable for multi-person recordings. A meeting transcript, for example, can preserve who made each statement; an interview can distinguish the interviewer from the guest; and a customer-support recording can separate an agent's comments from a caller's responses.
The model is available through OpenAI's audio Transcription API with the model identifier gpt-4o-transcribe-diarize. It belongs to OpenAI's speech-to-text model lineup rather than its general-purpose text-generation models. For ordinary transcription without speaker separation, OpenAI also offers the related GPT-4o Transcribe.
How speaker diarization works
To receive speaker-separated results, an application uses the diarized_json response format. The response contains transcript segments associated with speaker labels and timing metadata. An application can use those fields to display a readable dialogue, build a meeting summary grouped by participant, search what a particular speaker said, or export a structured conversation record.
The labels identify detected speakers, but they do not necessarily provide their real names automatically. OpenAI supports optional reference clips for known speakers to help map detected voices to known names. Up to four short reference clips can be supplied, and each clip should be between two and ten seconds long. This is useful when an application already knows the likely participants, such as a scheduled interview with a host and guest.
Reference clips should be treated as an aid to speaker identification rather than an absolute guarantee. Recordings with background noise, overlapping speech, poor microphone quality, or substantial changes in a speaker's voice may still require review. The supplied research does not provide an accuracy benchmark, so production teams should validate transcripts against representative recordings before relying on them without human checking.
Inputs, outputs and response formats
GPT-4o Transcribe Diarize accepts audio and produces text. Supported audio file formats include FLAC, MP3, MP4, MPEG, MPGA, M4A, OGG, WAV, and WebM. Its output is not synthesized speech: the model does not generate audio, music, or voice responses.
The documented response choices include regular JSON, plain text, and diarized JSON. The diarized format is the important option for this model's main use case because it exposes speaker-segment information and timing. The model can also stream transcription results as a recording is processed. Streaming can provide transcript events as processing progresses, while diarized responses can include finalized speaker-segment events.
For live microphone or media-stream transcription, OpenAI recommends separate realtime transcription models rather than treating this file-transcription model as a complete live-conversation solution. A related alternative in the available catalog is GPT-Live-Transcribe, which is positioned for realtime speech-to-text use cases.
Workflow and file-length limits
For recordings longer than 30 seconds, OpenAI requires a chunking strategy. The documented options include auto or a voice-activity-detection configuration. Chunking divides a longer recording into manageable sections, while voice activity detection can help identify portions containing speech. Applications should plan how to preserve ordering and combine segment results when processing a long meeting or podcast.
The model's documented context window is 16,000 tokens, and its maximum output is 2,000 tokens. These are model specifications rather than a promise that every audio file can be transcribed in one request. Long recordings may need to be divided into chunks, and the resulting transcript may need to be assembled by the application.
Word-level timestamp configuration is not available for this model's diarized response path. The supported diarized output instead provides segment-level start and end times. That distinction matters for workflows such as subtitle alignment, search-term highlighting, or applications that require precise timing for every word.
Main capabilities and limitations
- Audio input: The model processes recorded audio files in supported formats.
- Speaker separation: It identifies and labels different speakers within a recording.
- Segment timing: Diarized output includes segment start and end times.
- Known-speaker references: Up to four reference clips can help associate detected voices with known participants.
- Streaming: Transcription results can be streamed while a completed recording is processed.
- No image or video understanding: The model does not accept image or video input.
- No generated audio: It does not provide speech synthesis, music generation, or other audio output.
- No general text generation: Its primary output is a transcript, not a general-purpose written response.
- No prompting: Prompting is not supported for this diarization model.
- No word-level timestamps: The diarized response path provides segment timing rather than configurable word-by-word timing.
- No listed tool or function support: The supplied specifications identify no tool-use capability for this model.
These limitations make the model narrower than a general multimodal or language model. Its strength is the specific combination of transcription and speaker labeling, not broad reasoning, coding, web search, image analysis, or content generation.
Pricing and model specifications
| Specification | GPT-4o Transcribe Diarize |
|---|---|
| Provider | OpenAI |
| Model type | Automatic speech recognition with speaker diarization |
| Input | Audio |
| Output | Text, including diarized JSON responses |
| Context length | 16,000 tokens |
| Maximum output | 2,000 tokens |
| Input price | $2.50 per 1 million audio tokens |
| Output price | $10.00 per 1 million audio tokens |
| Status | Deprecated; currently accessible according to the supplied research |
The listed price is token-based rather than a simple per-minute rate. Actual usage depends on how audio is represented and processed, so teams should estimate costs from representative files and include any additional processing required to chunk and combine long recordings. The supplied research does not provide a separate free tier, minimum charge, or fixed per-hour conversion.
Speed, cost and capability trade-offs
The supplied editorial scoring rates the model's speed at 8 out of 10 and its cost at 7 out of 10. These are evaluation scores, not OpenAI-published benchmarks or guarantees. They indicate that the model is viewed as a relatively efficient choice for its specialized task, but they should not be interpreted as a measured latency or universal cost ranking.
Compared with a basic speech-to-text workflow, GPT-4o Transcribe Diarize adds speaker identification and segment timing, which can reduce the amount of custom post-processing needed for interviews and meetings. That additional structure is the reason to choose it over ordinary transcription. However, if an application only needs a single block of text, the diarization features may add unnecessary complexity.
Compared with realtime transcription options, this model is better suited to processing recorded files through the Transcription API. A realtime model may be more appropriate when text must appear continuously during a live conversation or media stream. Compared with a general-purpose language model, GPT-4o Transcribe Diarize is the more direct tool for producing a speaker-labeled transcript, but it is not the appropriate choice for broad reasoning, coding, web research, or post-transcription analysis by itself.
Best use cases
- Meeting records: Preserve statements by participant before generating minutes or action items.
- Interviews: Separate interviewer questions from guest answers and retain timing for review.
- Customer-support calls: Distinguish agent and customer speech for quality review or searchable records.
- Podcasts: Create speaker-labeled transcripts for editing, accessibility, or content repurposing.
- Research interviews: Organize responses by participant while retaining segment timestamps.
- Known-participant recordings: Use short reference clips when an application needs to map detected speakers to expected names.
For sensitive recordings, organizations should also consider their own retention, access-control, consent, and privacy requirements. The supplied model research describes technical behavior and pricing but does not establish a specific compliance certification or retention guarantee for this model.
When to choose this model
Choose GPT-4o Transcribe Diarize when the recording contains multiple speakers and the speaker identity of each segment is a core requirement. It is especially useful when an application would otherwise need to combine ordinary transcription with a separate diarization system. The model is also a reasonable fit when segment-level timestamps and optional known-speaker references are more valuable than word-level timing.
Choose a regular transcription model instead when speaker labels are unnecessary, when the workflow is optimized for simpler text output, or when the application needs capabilities outside speech recognition. Choose a realtime transcription option when audio arrives continuously and the user needs an immediate transcript. Choose a separate analysis or language model after transcription when the next task is summarization, classification, extraction, coding, or question answering.
New deployments should consider the model's lifecycle status before committing to it. The supplied research says that GPT-4o Transcribe Diarize was deprecated on August 26, 2026, remains accessible as of September 23, 2026, and is scheduled for API removal on February 26, 2027. OpenAI identifies GPT Live Transcribe or GPT Transcribe as migration targets. Therefore, this model may still be useful for an existing compatible workflow or short-term evaluation, but it is a poor long-term default for a new system unless the team has a clear migration plan.
Bottom line
GPT-4o Transcribe Diarize is a focused audio transcription model whose defining feature is speaker-labeled output. It supports common audio formats, diarized JSON, segment timestamps, streaming events, and up to four known-speaker reference clips. Its 16,000-token context window, 2,000-token maximum output, and token-based pricing are useful planning constraints, while the lack of prompting and word-level timestamps limits some advanced workflows.
Its strongest reason for selection is straightforward: it can turn a multi-speaker recording into a structured transcript without requiring speaker separation to be built as a completely separate step. Its most important reason for caution is equally clear: the model is deprecated and scheduled for removal, so new production systems should evaluate the named replacement options and design for migration.

