GPT-4o Mini

GPT-4o Mini Transcribe

by OpenAI · Current and available

OpenAI's lower-cost GPT-4o-based speech-to-text model for multilingual audio transcription, streamed file processing, prompting with domain vocabulary, and realtime transcription workflows.

Text Reasoning Coding
GPT-4o Mini Transcribe is OpenAI's specialized speech-recognition model for turning audio into text. It is available through the Audio API for uploaded recordings, streamed file transcription, and realtime transcription sessions. The model is designed to offer more accurate transcription and language recognition than the original Whisper models while costing less than GPT-4o Transcribe.
Outputs

What GPT-4o Mini Transcribe can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-4o Mini
Model type Other
Context window 16K tokens
Maximum output 2K tokens
Knowledge cutoff 2024-06-01
Release date 2025-03-20
Status Current and available
Knowledge cutoff notes

The official model page lists June 1, 2024 as the model knowledge cutoff. This is separate from the date of the audio being transcribed and is not extended by realtime processing or external application context.

Model notes

Specialized speech-to-text model powered by GPT-4o Mini. The model accepts audio input and returns text. OpenAI lists a 16,000-token context window and 2,000 maximum output tokens. Current pricing is based on audio tokens: $1.25 per 1 million input tokens and $5.00 per 1 million output tokens. The file-transcription workflow supports common formats including flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, and webm, with a 25 MB file limit documented for uploaded recordings. Prompting is supported for domain vocabulary, names, acronyms, punctuation, capitalization, and recording context. File transcription can stream partial results, and the model is also available for realtime transcription sessions. OpenAI lists the canonical alias gpt-4o-mini-transcribe and dated snapshots gpt-4o-mini-transcribe-2025-03-20 and gpt-4o-mini-transcribe-2025-12-15. The reasoning and coding scores are editorial assessments of applicability, not vendor benchmarks.

Cost

Model pricing

Input $1.25 per 1 million audio tokens
Output $5.00 per 1 million audio tokens
Model guide

GPT-4o Mini Transcribe: Pricing, Features, Limits and API Support

GPT-4o Mini Transcribe is OpenAI's lower-cost GPT-4o-based speech-to-text model. It converts recorded or realtime audio into text, supports multilingual recognition, prompting for specialized vocabulary, streamed transcription, and realtime transcription sessions. Its main trade-off is specialization: it is cheaper than GPT-4o Transcribe and suited to audio-to-text pipelines, but it is not a general-purpose reasoning, coding, tool-calling, or content-generation model.

What is GPT-4o Mini Transcribe?

GPT-4o Mini Transcribe is a speech-to-text model from OpenAI, powered by GPT-4o Mini. It accepts audio and returns a text transcript rather than generating audio, images, video, or general-purpose answers. OpenAI introduced it on March 20, 2025, alongside GPT-4o Transcribe and GPT-4o Mini TTS.

The model is intended for applications that need to understand spoken content in recordings or live audio streams. Typical examples include meeting transcription, customer-support call processing, voice-note conversion, accessibility features, searchable audio archives, and pipelines that turn speech into notes or other text processed by a separate application.

OpenAI positions the model as a lower-cost alternative to GPT-4o Transcribe. The provider also states that its transcription accuracy and language recognition are improved compared with the original Whisper models. Those are provider positioning claims; the supplied research does not provide an independent benchmark or a numerical accuracy comparison.

Capabilities and supported modalities

GPT-4o Mini Transcribe has one central capability: converting spoken audio into text. Its modality profile is therefore straightforward:

  • Audio input: supported.
  • Text input: supported for transcription context and prompting.
  • Text output: supported.
  • Audio, image, video, or music output: not supported.

The model supports multilingual speech recognition and language identification. Prompting lets an application provide context that may help with names, acronyms, product terms, technical vocabulary, punctuation, capitalization, or other words likely to be difficult for a transcription system. For example, a customer-support application could provide the names of products used by its company before transcribing a call.

Its listed context window is 16,000 tokens, and its maximum output length is 2,000 tokens. These are model specifications supplied by OpenAI. They should not be interpreted as a promise that every audio recording will produce a transcript of exactly that size: the practical result depends on the audio, speech rate, language, and transcription workflow.

Recorded and realtime transcription workflows

For a completed recording, an application can upload an audio file to the Audio Transcriptions workflow and receive a transcript. The documented file formats include FLAC, MP3, MP4, MPEG, MPGA, M4A, OGG, WAV, and WebM. The uploaded-file workflow has a documented 25 MB limit.

File transcription can also stream partial text while a completed recording is being processed. Streaming is useful when a user should see a transcript develop without waiting for the entire response, although partial text may be incomplete until later events finish the relevant segment.

For continuously arriving audio, such as microphone input, a phone call, or a media stream, the realtime transcription path is more appropriate than repeatedly uploading files. Realtime sessions are designed for audio that arrives over time and can return transcript events as processing progresses. The model supports prompting in this workflow as well, allowing an application to supply recording-specific vocabulary or context.

These workflows solve different operational problems. File transcription is a natural fit for recordings that already exist, while realtime transcription is intended for live or ongoing audio. The research confirms support for both paths but does not specify a universal latency guarantee.

Pricing and model versions

The listed price is $1.25 per 1 million audio input tokens and $5.00 per 1 million audio output tokens. This is token-based audio pricing, not a flat per-minute price. Actual cost depends on how much audio is processed and how many output tokens the transcript produces.

OpenAI lists gpt-4o-mini-transcribe as the canonical model alias. Dated snapshots include gpt-4o-mini-transcribe-2025-03-20 and gpt-4o-mini-transcribe-2025-12-15. The alias is generally convenient when an application should follow the current model identity, while a dated snapshot can be useful when a team needs to lock its integration to a particular version. Snapshot behavior and continued availability should be checked in OpenAI's current documentation before production deployment.

The model is available through OpenAI's API. The supplied research does not identify a consumer ChatGPT interface that exposes GPT-4o Mini Transcribe as a selectable standalone model, so it is best evaluated as an API component for audio-processing applications.

Main strengths and trade-offs

Lower-cost specialized transcription

Its clearest advantage is the balance between transcription capability and price. Compared with GPT-4o Transcribe, GPT-4o Mini Transcribe is positioned as the more economical choice when an application needs text transcripts rather than the broader capabilities of a larger audio model. This makes it relevant for high-volume workloads such as call archives, meeting recordings, voice notes, and searchable media.

Language recognition and domain context

Multilingual recognition and language identification make the model suitable for applications that process more than one spoken language. Prompting can provide terminology that may otherwise be difficult to recognize, including proper names, acronyms, and industry-specific words. Prompting is helpful context, not a guarantee of error-free transcription.

Streaming and realtime support

The ability to stream partial transcript results from file processing and use realtime transcription sessions gives developers more than a simple upload-and-wait workflow. A live captioning or note-taking interface can begin displaying text while audio continues to arrive or while processing is still underway.

Specialization is also a limitation

GPT-4o Mini Transcribe is not a general-purpose language model for reasoning, coding, web search, or arbitrary text generation. It does not natively generate audio, images, or video. The research also lists tool use, function calling, structured output, fine-tuning, caching, and batch API support as unavailable for this model. Applications that need those capabilities must use another model or add a separate processing stage.

Speaker identification is another important boundary. The supplied research says applications needing speaker identification should evaluate OpenAI's dedicated diarization model instead. GPT-4o Mini Transcribe should therefore not be selected solely because an application needs to determine which person said each sentence.

Reasoning, coding, and tool support

This model's job is transcription, not reasoning over a problem or writing software. The research gives it editorial reasoning and coding applicability scores of 1, but those scores are assessments rather than OpenAI benchmarks. In practical terms, the model should be treated as having no intended general-purpose reasoning or coding role.

It also should not be treated as a tool-using assistant. The supplied model record marks tool use as unsupported and web search as unsupported. A common architecture is to use GPT-4o Mini Transcribe first, then send the resulting text to another model or application component for summarization, extraction, search, classification, or workflow actions. That division keeps transcription and downstream interpretation separate.

When to choose GPT-4o Mini Transcribe

Choose GPT-4o Mini Transcribe when the central requirement is affordable audio-to-text conversion and the application can work with text output. It is a strong candidate for:

  • meeting and interview transcription;
  • customer-support and call-center recordings;
  • voice-note and voicemail conversion;
  • multilingual audio transcription;
  • live captions or transcript displays using realtime audio;
  • audio search pipelines that index spoken content; and
  • applications that need prompting for specialized terms, names, or acronyms.

It is particularly sensible when cost and throughput matter more than access to a broad conversational model. The lower listed price than GPT-4o Transcribe can make it a better fit for routine transcription at scale, provided its accuracy is acceptable for the application's language, audio quality, and terminology.

When another option may be better

Choose GPT-4o Transcribe when the application can justify a higher transcription cost in exchange for the capabilities or quality level associated with that model. The supplied research establishes the pricing and positioning difference, but it does not provide a numerical quality threshold, so teams should test representative recordings before deciding.

Choose a dedicated diarization solution when the transcript must identify or separate speakers. Choose a general-purpose model after transcription when the workflow requires reasoning, coding, structured extraction, web search, or tool calls. A separate model may also be more suitable when the desired result is generated speech, an image, video, or another non-text output.

The original Whisper models may still be relevant in environments where an existing Whisper-based deployment is preferred, but the supplied research does not provide a current price, capability matrix, or deployment comparison for them. GPT-4o Mini Transcribe is the more directly supported choice here when an application wants OpenAI's current Audio API transcription workflow.

Bottom line

GPT-4o Mini Transcribe is a focused, lower-cost OpenAI model for converting recorded or realtime audio into text. Its verified specifications include audio input, text output, a 16,000-token context window, a 2,000-token maximum output, streaming support, realtime transcription support, and pricing of $1.25 per 1 million audio input tokens plus $5.00 per 1 million audio output tokens.

Its value comes from doing one job efficiently rather than from covering every AI task. It is a practical choice for transcription pipelines that need multilingual recognition, domain prompting, and economical processing. It is not the right standalone choice for speaker diarization, general reasoning, coding, tool use, or content generation, so those requirements should be handled by other components.


Answers to Frequently Asked Questions

What are the main limitations of GPT-4o Mini Transcribe?
GPT-4o Mini Transcribe is specialized for audio-to-text conversion and is not a general-purpose model for reasoning, coding, web search, tool use, structured output, or content generation. It does not generate audio, images, or video, and applications that need speaker identification should use a dedicated diarization solution.
Does GPT-4o Mini Transcribe support realtime transcription and streaming?
Yes. GPT-4o Mini Transcribe supports transcription of completed audio files, streaming partial transcript results during file processing, and realtime transcription sessions for continuously arriving audio such as microphone input, phone calls, and media streams.
What are the file size, context, and output limits for GPT-4o Mini Transcribe?
The Audio Transcriptions file-upload workflow has a documented 25 MB file limit. The model has a 16,000-token context window and a maximum output length of 2,000 tokens. Practical transcript length depends on factors such as recording duration, speech rate, language, and workflow.
What is GPT-4o Mini Transcribe used for?
GPT-4o Mini Transcribe is an OpenAI speech-to-text model for converting recorded or realtime audio into text. Common uses include meeting transcription, customer-support call processing, voice-note conversion, live captions, accessibility features, and searchable audio archives.
How much does GPT-4o Mini Transcribe cost?
GPT-4o Mini Transcribe costs $1.25 per 1 million audio input tokens and $5.00 per 1 million audio output tokens. Pricing is token-based rather than a flat per-minute rate, so the total cost depends on the audio processed and the transcript length.


Sources 5
Provider

About OpenAI