What is Whisper?
Whisper is OpenAI's general-purpose automatic speech recognition model. Automatic speech recognition, or ASR, means converting spoken audio into written text. Whisper can also identify the language being spoken and translate speech from other languages into English.
The name covers two related things. OpenAI's original Whisper project is an open-source speech recognition system with downloadable model checkpoints in several sizes. Separately, whisper-1 is the hosted model identifier available through the OpenAI API. These should not be treated as identical products: the API model is a managed service, while the open-source project can be run independently using its published checkpoints.
Whisper was introduced by OpenAI on September 21, 2022. Its main role in OpenAI's catalog is speech-to-text rather than general conversation, reasoning, coding, image understanding, or speech generation.
Current status and position in OpenAI's catalog
The hosted whisper-1 API model is deprecated but remains accessible during its transition period. OpenAI's supplied documentation lists a shutdown date of February 26, 2027, and recommends migration to newer speech-oriented options such as gpt-live-transcribe or gpt-transcribe.
This status matters when choosing Whisper for a new application. It may still be useful for an existing integration, a short-lived project, or a workflow whose requirements match its simple transcription interface. However, a system expected to operate beyond the shutdown date should not make Whisper its long-term dependency without a documented migration plan.
What Whisper can do
- Audio transcription: converts spoken audio into text.
- Speech translation: translates non-English speech into English text.
- Language identification: identifies the language present in the audio.
- Timestamped output: provides segment- or word-level timing where supported, which is useful for captions and searchable media.
- Multiple output formats: returns results as JSON, plain text, SRT, verbose JSON, and VTT.
OpenAI documentation identifies support for 98 languages. The practical quality of transcription can vary with language, recording conditions, accents, background noise, overlapping speech, and the subject matter of the recording. The supplied specifications do not provide a universal accuracy score, so transcription should be reviewed when errors would affect legal, medical, financial, compliance, or publication decisions.
Inputs, outputs and technical limits
Whisper accepts audio input and produces text output. It does not accept images or video as model inputs in the documented configuration, and it does not directly produce audio, images, or video.
| Capability | Whisper specification |
|---|---|
| Hosted model ID | whisper-1 |
| Primary input | Audio |
| Primary output | Text |
| Supported languages | 98 languages according to OpenAI documentation |
| Translation direction | Speech in other languages translated into English |
| Timestamp formats | Segment- or word-level timestamps where supported |
| Output formats | JSON, plain text, SRT, verbose JSON, and VTT |
| Published context length | Not specified in the supplied documentation |
| Published maximum output tokens | Not specified in the supplied documentation |
A token limit is not the most useful way to evaluate this model because Whisper is designed around audio transcription rather than open-ended text generation. Nevertheless, OpenAI's supplied model information does not state a context-window or maximum-output-token figure, so applications should follow the current API documentation for accepted audio-file constraints and response behavior rather than assuming a limit.
Whisper API pricing
The supplied OpenAI pricing information lists Whisper at $0.006 per minute of audio. This is an audio-duration price, not a recurring monthly subscription. The final cost depends on the amount of audio processed, so a one-hour recording would be priced according to 60 minutes of submitted audio before any applicable account or platform considerations.
There is no separate output-token charge listed for Whisper in the supplied research. The model is therefore straightforward to estimate for batch transcription: multiply the number of processed audio minutes by $0.006. Teams should still verify the live pricing page before deployment because provider prices and model availability can change, particularly while the model is being retired.
Speed, cost and capability trade-offs
Whisper's strongest trade-off is focused functionality. It is less expensive and conceptually simpler than using a general-purpose multimodal or real-time model when the task is simply turning recorded speech into text. It also offers useful subtitle-oriented formats and timestamps without requiring a general conversational model.
Its limitations are equally important. Whisper is not documented here as a reasoning model, coding model, tool-using agent, speaker-diarization system, or speech-generation model. It does not provide native image or video understanding, and it is not intended to answer questions about the transcript by itself. A workflow that needs summarization, extraction, classification, or interactive conversation may need a separate text model after transcription or a newer model that combines speech processing with additional capabilities.
The supplied editorial assessment rates its speed and cost favorably, but those are editorial scores rather than OpenAI-published benchmarks. No benchmark result or guaranteed latency figure is provided, so they should not be interpreted as service-level guarantees.
Main strengths and limitations
Strengths
- Broad multilingual coverage: useful for transcribing content in many languages and translating speech into English.
- Practical captioning support: SRT, VTT, verbose JSON, and timestamp capabilities fit subtitle and media workflows.
- Predictable pricing: the published rate is based on audio minutes, making batch-processing estimates relatively simple.
- Focused interface: teams that only need speech recognition do not need to use a general-purpose conversational model for the initial transcription step.
- Open-source lineage: the related Whisper project can be studied or run separately from the hosted API, subject to its own deployment requirements.
Limitations
- Deprecation: the hosted API model is scheduled for shutdown on February 26, 2027.
- No documented context or output-token figures: the supplied research does not state these limits.
- No speaker diarization: the research specifically lists speaker diarization as outside its intended use.
- Text output only: it does not generate natural speech or other media.
- Not a reasoning system: it transcribes and translates audio but is not presented as a model for analysis, coding, or general text generation.
- Transcription is not guaranteed to be error-free: recordings with noise, accents, multiple speakers, or difficult terminology should be checked.
Best use cases for Whisper
Whisper is a reasonable fit when the central requirement is converting recorded audio into usable text at a known per-minute cost. Examples include:
- Creating draft transcripts of interviews, meetings, lectures, podcasts, and recordings.
- Generating subtitle files in SRT or VTT format.
- Building searchable text archives from multilingual audio.
- Translating non-English speech into English text for review or downstream processing.
- Adding timestamps to media transcripts for navigation and caption editing.
- Running a transcription stage before sending text to a separate summarization, search, or extraction system.
For a production pipeline, it is sensible to preserve the original audio, store the raw response, retain language and timestamp metadata, and include a review path for important transcripts. Because the hosted model has a published shutdown date, the pipeline should also isolate the transcription provider behind an interface that makes later migration easier.
When to choose Whisper
Choose Whisper when you need affordable, multilingual transcription or English translation of recorded audio, especially when subtitle formats and timestamps are important. It can be a practical choice for an existing system that already uses whisper-1 and has a near-term need for stable, focused speech recognition.
Choose another option when long-term availability is a requirement. OpenAI's supplied research specifically recommends considering gpt-live-transcribe or gpt-transcribe for migration. A real-time speech model may be more appropriate for live conversations or low-latency interaction, while a broader multimodal or general-purpose model may be preferable when the same system must interpret images, reason over transcripts, call tools, or generate conversational responses. Those alternatives may offer broader capabilities, but they can also introduce different pricing, latency, integration, and availability trade-offs.
Whisper is also a poor fit if the required output is a spoken response, if reliable speaker labels are essential, or if the application expects the model to perform complex reasoning directly over audio. In those cases, use a workflow or model designed for the missing capability rather than treating transcription as a complete conversational system.
Bottom line
Whisper remains a focused OpenAI speech recognition model for multilingual transcription, English speech translation, language identification, and timestamped text output. Its $0.006-per-minute pricing and familiar subtitle formats make it useful for recorded-audio workflows. The decisive drawback is its lifecycle: the hosted whisper-1 API is deprecated and scheduled to shut down on February 26, 2027. It can still serve compatible short-term or existing workloads, but new systems should evaluate the recommended successor models and design for migration from the outset.

