Whisper

Whisper

by OpenAI · Deprecated; currently accessible through the OpenAI API and scheduled for shutdown on February 26, 2027.

OpenAI Whisper is a multilingual automatic speech recognition model that converts audio to text, translates speech into English, identifies languages, and supports timestamped transcription in formats such as SRT and VTT. The whisper-1 API model costs $0.006 per minute but is deprecated and scheduled for shutdown on February 26, 2027.

Text
Whisper is an OpenAI speech recognition model trained on a large and diverse multilingual audio dataset. It can transcribe speech into text, translate non-English speech into English, identify languages, and generate timestamped output for subtitles, captions, and media workflows. The hosted whisper-1 API model is currently available but deprecated, so new production systems should consider a supported successor instead.
Outputs

What Whisper can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Whisper
Model type Other
Release date September 21, 2022
Status Deprecated; currently accessible through the OpenAI API and scheduled for shutdown on February 26, 2027.
Deprecation date 2026-08-26
Shutdown date 2027-02-26
Knowledge cutoff notes

OpenAI does not publish a model knowledge cutoff for Whisper. As a speech recognition model, its relevant behavior is determined by its trained audio and language representations rather than a documented text knowledge-cutoff date.

Model notes

The canonical hosted API identifier is whisper-1, powered by OpenAI's open-source Whisper V2 model. It accepts audio input and returns text in formats including JSON, plain text, SRT, verbose JSON, and VTT. It supports transcription in 98 languages according to OpenAI's documentation and can translate non-English speech into English. Timestamp granularities are supported specifically for whisper-1. The underlying open-source Whisper project includes multiple downloadable model sizes and variants, but those checkpoints are distinct from the hosted whisper-1 API alias. OpenAI announced deprecation on August 26, 2026 and recommends migrating to gpt-live-transcribe or gpt-transcribe.

Cost

Model pricing

Input $0.006 per minute of audio
Model guide

Whisper: OpenAI Speech Recognition, Transcription, Pricing and API Status

Whisper is OpenAI's multilingual automatic speech recognition model for converting audio to text, translating speech into English, identifying languages, and producing timestamped transcripts. Its hosted whisper-1 API model remains accessible but is deprecated and scheduled for shutdown on February 26, 2027.

What is Whisper?

Whisper is OpenAI's general-purpose automatic speech recognition model. Automatic speech recognition, or ASR, means converting spoken audio into written text. Whisper can also identify the language being spoken and translate speech from other languages into English.

The name covers two related things. OpenAI's original Whisper project is an open-source speech recognition system with downloadable model checkpoints in several sizes. Separately, whisper-1 is the hosted model identifier available through the OpenAI API. These should not be treated as identical products: the API model is a managed service, while the open-source project can be run independently using its published checkpoints.

Whisper was introduced by OpenAI on September 21, 2022. Its main role in OpenAI's catalog is speech-to-text rather than general conversation, reasoning, coding, image understanding, or speech generation.

Current status and position in OpenAI's catalog

The hosted whisper-1 API model is deprecated but remains accessible during its transition period. OpenAI's supplied documentation lists a shutdown date of February 26, 2027, and recommends migration to newer speech-oriented options such as gpt-live-transcribe or gpt-transcribe.

This status matters when choosing Whisper for a new application. It may still be useful for an existing integration, a short-lived project, or a workflow whose requirements match its simple transcription interface. However, a system expected to operate beyond the shutdown date should not make Whisper its long-term dependency without a documented migration plan.

What Whisper can do

  • Audio transcription: converts spoken audio into text.
  • Speech translation: translates non-English speech into English text.
  • Language identification: identifies the language present in the audio.
  • Timestamped output: provides segment- or word-level timing where supported, which is useful for captions and searchable media.
  • Multiple output formats: returns results as JSON, plain text, SRT, verbose JSON, and VTT.

OpenAI documentation identifies support for 98 languages. The practical quality of transcription can vary with language, recording conditions, accents, background noise, overlapping speech, and the subject matter of the recording. The supplied specifications do not provide a universal accuracy score, so transcription should be reviewed when errors would affect legal, medical, financial, compliance, or publication decisions.

Inputs, outputs and technical limits

Whisper accepts audio input and produces text output. It does not accept images or video as model inputs in the documented configuration, and it does not directly produce audio, images, or video.

CapabilityWhisper specification
Hosted model IDwhisper-1
Primary inputAudio
Primary outputText
Supported languages98 languages according to OpenAI documentation
Translation directionSpeech in other languages translated into English
Timestamp formatsSegment- or word-level timestamps where supported
Output formatsJSON, plain text, SRT, verbose JSON, and VTT
Published context lengthNot specified in the supplied documentation
Published maximum output tokensNot specified in the supplied documentation

A token limit is not the most useful way to evaluate this model because Whisper is designed around audio transcription rather than open-ended text generation. Nevertheless, OpenAI's supplied model information does not state a context-window or maximum-output-token figure, so applications should follow the current API documentation for accepted audio-file constraints and response behavior rather than assuming a limit.

Whisper API pricing

The supplied OpenAI pricing information lists Whisper at $0.006 per minute of audio. This is an audio-duration price, not a recurring monthly subscription. The final cost depends on the amount of audio processed, so a one-hour recording would be priced according to 60 minutes of submitted audio before any applicable account or platform considerations.

There is no separate output-token charge listed for Whisper in the supplied research. The model is therefore straightforward to estimate for batch transcription: multiply the number of processed audio minutes by $0.006. Teams should still verify the live pricing page before deployment because provider prices and model availability can change, particularly while the model is being retired.

Speed, cost and capability trade-offs

Whisper's strongest trade-off is focused functionality. It is less expensive and conceptually simpler than using a general-purpose multimodal or real-time model when the task is simply turning recorded speech into text. It also offers useful subtitle-oriented formats and timestamps without requiring a general conversational model.

Its limitations are equally important. Whisper is not documented here as a reasoning model, coding model, tool-using agent, speaker-diarization system, or speech-generation model. It does not provide native image or video understanding, and it is not intended to answer questions about the transcript by itself. A workflow that needs summarization, extraction, classification, or interactive conversation may need a separate text model after transcription or a newer model that combines speech processing with additional capabilities.

The supplied editorial assessment rates its speed and cost favorably, but those are editorial scores rather than OpenAI-published benchmarks. No benchmark result or guaranteed latency figure is provided, so they should not be interpreted as service-level guarantees.

Main strengths and limitations

Strengths

  • Broad multilingual coverage: useful for transcribing content in many languages and translating speech into English.
  • Practical captioning support: SRT, VTT, verbose JSON, and timestamp capabilities fit subtitle and media workflows.
  • Predictable pricing: the published rate is based on audio minutes, making batch-processing estimates relatively simple.
  • Focused interface: teams that only need speech recognition do not need to use a general-purpose conversational model for the initial transcription step.
  • Open-source lineage: the related Whisper project can be studied or run separately from the hosted API, subject to its own deployment requirements.

Limitations

  • Deprecation: the hosted API model is scheduled for shutdown on February 26, 2027.
  • No documented context or output-token figures: the supplied research does not state these limits.
  • No speaker diarization: the research specifically lists speaker diarization as outside its intended use.
  • Text output only: it does not generate natural speech or other media.
  • Not a reasoning system: it transcribes and translates audio but is not presented as a model for analysis, coding, or general text generation.
  • Transcription is not guaranteed to be error-free: recordings with noise, accents, multiple speakers, or difficult terminology should be checked.

Best use cases for Whisper

Whisper is a reasonable fit when the central requirement is converting recorded audio into usable text at a known per-minute cost. Examples include:

  • Creating draft transcripts of interviews, meetings, lectures, podcasts, and recordings.
  • Generating subtitle files in SRT or VTT format.
  • Building searchable text archives from multilingual audio.
  • Translating non-English speech into English text for review or downstream processing.
  • Adding timestamps to media transcripts for navigation and caption editing.
  • Running a transcription stage before sending text to a separate summarization, search, or extraction system.

For a production pipeline, it is sensible to preserve the original audio, store the raw response, retain language and timestamp metadata, and include a review path for important transcripts. Because the hosted model has a published shutdown date, the pipeline should also isolate the transcription provider behind an interface that makes later migration easier.

When to choose Whisper

Choose Whisper when you need affordable, multilingual transcription or English translation of recorded audio, especially when subtitle formats and timestamps are important. It can be a practical choice for an existing system that already uses whisper-1 and has a near-term need for stable, focused speech recognition.

Choose another option when long-term availability is a requirement. OpenAI's supplied research specifically recommends considering gpt-live-transcribe or gpt-transcribe for migration. A real-time speech model may be more appropriate for live conversations or low-latency interaction, while a broader multimodal or general-purpose model may be preferable when the same system must interpret images, reason over transcripts, call tools, or generate conversational responses. Those alternatives may offer broader capabilities, but they can also introduce different pricing, latency, integration, and availability trade-offs.

Whisper is also a poor fit if the required output is a spoken response, if reliable speaker labels are essential, or if the application expects the model to perform complex reasoning directly over audio. In those cases, use a workflow or model designed for the missing capability rather than treating transcription as a complete conversational system.

Bottom line

Whisper remains a focused OpenAI speech recognition model for multilingual transcription, English speech translation, language identification, and timestamped text output. Its $0.006-per-minute pricing and familiar subtitle formats make it useful for recorded-audio workflows. The decisive drawback is its lifecycle: the hosted whisper-1 API is deprecated and scheduled to shut down on February 26, 2027. It can still serve compatible short-term or existing workloads, but new systems should evaluate the recommended successor models and design for migration from the outset.


Answers to Frequently Asked Questions

What are Whisper's main limitations?
Whisper does not directly provide reasoning, coding, speaker diarization, image or video understanding, or speech generation. Transcription quality can also vary with noise, accents, overlapping speakers, language, and specialized terminology, so important transcripts should be reviewed.
What audio and output formats does Whisper support?
Whisper accepts audio and returns text in formats including JSON, plain text, verbose JSON, SRT, and VTT. Segment- or word-level timestamps are available where supported.
Is the Whisper API still available?
The hosted whisper-1 API model is deprecated but remains accessible during its transition period. OpenAI lists a planned shutdown date of February 26, 2027, and recommends evaluating gpt-live-transcribe or gpt-transcribe for migration.
What is OpenAI Whisper used for?
Whisper is an automatic speech recognition model used to convert audio into text, identify spoken languages, translate non-English speech into English, and produce timestamped outputs for captions and searchable media.
How much does the Whisper API cost?
The supplied OpenAI pricing lists Whisper at $0.006 per minute of processed audio. For example, one hour of audio would cost approximately $0.36 before any applicable account or platform considerations.


Sources 6
Provider

About OpenAI