Cohere Transcribe

Cohere Transcribe

by Cohere · Live; open-source research release

Cohere Transcribe is a 2-billion-parameter Conformer-based speech recognition model that converts uploaded audio into text in 14 languages. It supports common audio formats up to 25 MB and is available through Cohere's Audio Transcriptions API and Model Vault. The model is designed for efficient transcription, but requires a specified input language and does not document automatic language detection, timestamps, speaker diarization, or live streaming.

Text Reasoning Coding
Cohere Transcribe is a dedicated speech-to-text model for converting audio recordings into written text. The approximately 2-billion-parameter model supports 14 languages, accepts common audio formats up to 25 MB, and is available through Cohere's Audio Transcriptions API and Model Vault. It is a focused automatic speech recognition system rather than a general-purpose text model: it produces transcripts, but it is not designed for text generation, speech synthesis, speaker separation, or real-time conversational audio.
Outputs

What Cohere Transcribe can produce

Text
Inputs

What it can understand

Audio
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Cohere Transcribe
Model type Speech Recognition
Release date 2026-03-26
Status Live; open-source research release
Knowledge cutoff notes

Cohere has not published a knowledge cutoff date for this speech recognition model. It is trained for audio-to-text transcription rather than general factual question answering.

Model notes

Cohere Transcribe is a 2B-parameter Conformer-based encoder-decoder ASR model. It supports English, German, French, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Vietnamese, Chinese, Arabic, Japanese, and Korean. The API requires the language to be specified and accepts FLAC, MP3, MPEG, MPGA, OGG, and WAV files up to 25 MB. Cohere describes the model as open source under Apache 2.0. API experimentation is free subject to rate limits. Production deployment without those limits is available through Model Vault, which is priced per hour-instance. Cohere states that the model can achieve a real-time factor up to three times faster than comparable dedicated ASR models in the same size range. Editorial scores are comparative estimates for the speech-recognition category, not provider benchmarks.

Cost

Model pricing

Input No per-input token price published. API experimentation is free subject to rate limits; Model Vault uses hourly instance-based pricing.
Output No per-output token price published. API experimentation is free subject to rate limits; Model Vault uses hourly instance-based pricing.
Model guide

Cohere Transcribe: Open-Source Multilingual Speech Recognition for Efficient Audio-to-Text

Cohere Transcribe is a 2-billion-parameter, open-source automatic speech recognition model from Cohere. It converts uploaded audio into text in 14 languages through Cohere's Audio Transcriptions API or Model Vault. Its main advantages are multilingual coverage, an efficient Conformer-based design, Apache 2.0 licensing, and deployment options for enterprise workloads. It requires the input language to be specified and does not provide documented automatic language detection, timestamps, speaker diarization, or live streaming transcription.

What is Cohere Transcribe?

Cohere Transcribe is Cohere's automatic speech recognition (ASR) model. In practical terms, it takes an audio file containing spoken language and returns the words as text. The canonical model identifier is cohere-transcribe-03-2026.

The model is an open-source research release with approximately 2 billion parameters and is distributed under the Apache 2.0 license. It is available through Cohere's Audio Transcriptions API for hosted experimentation and application integration, and through Cohere Model Vault for production deployments that require dedicated infrastructure. This places it in Cohere's catalog as a specialized audio model alongside, rather than as a replacement for, the company's text-generation, retrieval, embedding, and document-processing systems.

Unlike a general multimodal language model, Cohere Transcribe has a narrow purpose: converting speech into text. Its documented output is text, and it does not generate audio, images, video, or general-purpose written responses.

Supported languages and audio input

Cohere Transcribe supports 14 languages:

  • English
  • German
  • French
  • Italian
  • Spanish
  • Portuguese
  • Greek
  • Dutch
  • Polish
  • Vietnamese
  • Chinese
  • Arabic
  • Japanese
  • Korean

The transcription request requires the caller to provide the input language using an ISO-639-1 code. The model is optimized for one specified language at a time, and the supplied documentation does not describe explicit automatic language detection. Applications handling recordings in unknown or mixed languages therefore need to identify the language before transcription or add a separate language-identification step.

The documented API accepts FLAC, MP3, MPEG, MPGA, OGG, and WAV files. The maximum documented upload size is 25 MB. This makes the model suitable for ordinary meeting recordings, interviews, voice notes, and call segments, but larger recordings may need to be divided before submission. The supplied documentation describes an uploaded-file workflow rather than a live streaming transcription interface.

Architecture and performance positioning

Cohere describes Transcribe as using a speech-optimized Conformer-based encoder-decoder architecture. A Conformer combines techniques suited to speech processing with transformer-style sequence modeling. At a high level, the system converts the audio waveform into log-Mel spectrogram features, processes those features with a Conformer encoder, and uses a lightweight Transformer decoder to generate text tokens.

Cohere says the model was trained from scratch with supervised cross-entropy and optimized for low word error rate and efficient serving. These are provider descriptions of the model's design and goals, not an independent benchmark result. Cohere also states that Transcribe can achieve a real-time factor up to three times faster than other dedicated ASR models in the same size range. The claim indicates a focus on throughput and latency, but actual performance will depend on audio characteristics, hardware, deployment configuration, and workload.

The model's relatively focused architecture is important to its positioning. It is not intended to reason over a transcript, write a report, call tools, or answer questions about the recording in the same request. A workflow can transcribe audio first and then send the resulting text to another model or application for summarization, search, classification, or extraction.

API and deployment options

The Audio Transcriptions API provides the simplest documented way to test or integrate Cohere Transcribe. A caller uploads a supported audio file, specifies the language, and receives the transcribed text. Cohere provides free, low-setup API experimentation subject to rate limits. The research does not provide a per-minute, per-input-token, or per-output-token price for the hosted API.

For production use without the trial rate limits, Cohere makes the model available through Model Vault. Model Vault pricing is calculated per hour-instance rather than by input or output token. Cohere also offers longer-term commitment discounts, but the supplied research does not include a public numeric hourly price. Organizations should therefore request or confirm a deployment quote instead of estimating cost from token-based language-model pricing.

Model Vault is relevant when an organization needs a more controlled serving arrangement or predictable dedicated capacity. The hosted API is more appropriate for initial evaluation and applications that can work within its file and rate limits. Neither option is documented here as providing live streaming transcription, so real-time voice applications should verify current product support or consider a service specifically designed for streaming audio.

Capabilities and limitations

The following distinctions summarize what is documented for the model:

AreaDocumented behavior
Primary taskAutomatic speech recognition and audio-to-text transcription
InputAudio waveform files
OutputText
Languages14 specified languages
File formatsFLAC, MP3, MPEG, MPGA, OGG, and WAV
Maximum file size25 MB
Language selectionCaller must specify the language
StreamingNo documented live streaming interface
TimestampsNot provided in the supplied specification
Speaker diarizationNot provided in the supplied specification
LicenseApache 2.0

There is no documented context window or maximum output-token limit for Cohere Transcribe. The relevant practical limit is the 25 MB maximum upload size. The response is a transcript, so output length will depend on the duration and speech content of the submitted recording rather than on a published general-purpose language-model token limit.

The model has no documented tool or function-calling support, JSON-mode support, caching, batch API, or fine-tuning capability in the supplied research. It also does not provide reasoning or coding capabilities in the usual language-model sense. Those omissions are not necessarily defects for transcription, but they matter when comparing it with a general-purpose model that can process audio and then perform additional tasks directly.

Main strengths and trade-offs

Where Cohere Transcribe is strong

  • Focused speech recognition: The model is purpose-built for converting speech into text rather than sharing capacity with unrelated generation tasks.
  • Multilingual coverage: Fourteen supported languages cover many common enterprise and international transcription scenarios.
  • Efficient deployment: Cohere positions its Conformer-based design for low word error rate and fast inference, including a claimed real-time factor advantage over comparable dedicated ASR models.
  • Open licensing: The Apache 2.0 license can be useful for organizations evaluating open-weight deployment and integration options, subject to their own legal and operational review.
  • Enterprise deployment path: Model Vault provides a production route based on dedicated hourly instances rather than token billing.

Where it is limited

  • No automatic language detection: The application must provide the language code, which complicates mixed-language or unknown-language recordings.
  • No documented timestamps: Users needing word-level or segment-level timing may need an additional processing system.
  • No speaker diarization: The model does not identify which participant spoke each portion of a recording.
  • File-based workflow: The documented API accepts uploaded files, not a live audio stream.
  • 25 MB upload ceiling: Long recordings may need preprocessing or segmentation.
  • Limited task scope: It transcribes audio but does not summarize, reason over, search, or transform the transcript by itself.

Best use cases

Cohere Transcribe is a practical fit when the central requirement is multilingual transcription of uploaded recordings. Examples include meeting notes, interview archives, voice-note conversion, customer-service or call-center recordings, multilingual enterprise speech repositories, and internal applications that need a text representation before applying search or language analysis.

It is especially suitable when an organization values an open-source release, wants to evaluate a relatively efficient dedicated ASR model, or plans to use Cohere's enterprise deployment options. A typical pipeline might upload a recording, specify its known language, store the returned transcript, and then pass that text to a separate search, summarization, or document workflow.

When to choose Cohere Transcribe

Choose Cohere Transcribe when you need a dedicated speech-to-text component with support for the listed 14 languages, common uploaded audio formats, and a path from rate-limited API experimentation to Model Vault deployment. Its combination of focused architecture, Apache 2.0 licensing, and Cohere's enterprise infrastructure options may be more relevant than a general-purpose multimodal model when transcription throughput and deployment control are the priorities.

Another speech service may be more appropriate if the application requires automatic language identification, live streaming, timestamps, or speaker separation. A general-purpose language model may be a better second-stage choice when the main task is to analyze the content, produce a structured report, answer questions, or generate code after transcription. Cohere Transcribe can serve as the audio-to-text stage in that workflow, but the supplied specifications do not indicate that it performs those later tasks itself.

Bottom line

Cohere Transcribe is a specialized, open-source multilingual ASR model rather than a broad conversational AI system. Its documented strengths are 14-language coverage, support for common audio files up to 25 MB, an efficient Conformer-based design, and access through both Cohere's transcription API and Model Vault. Its main constraints are equally clear: the language must be specified, the workflow is file-based, and timestamps, diarization, automatic language detection, and general reasoning are not documented. For enterprise audio-to-text workloads that fit those boundaries, it offers a focused alternative to using a larger general-purpose model for transcription.


Answers to Frequently Asked Questions

How can Cohere Transcribe be deployed?
Cohere Transcribe can be tested and integrated through Cohere's Audio Transcriptions API, which is subject to rate limits, or deployed for production through Cohere Model Vault using dedicated hourly instances. The model is released under the Apache 2.0 license.
Does Cohere Transcribe provide timestamps or speaker diarization?
Timestamps and speaker diarization are not provided in the supplied specification. Applications that need timing information or speaker identification may require additional processing tools.
Does Cohere Transcribe automatically detect languages or support live streaming?
The caller must provide the input language using an ISO-639-1 code, and automatic language detection is not documented. The documented workflow is based on uploaded files, with no live streaming transcription interface specified.
Which languages and audio formats does Cohere Transcribe support?
Cohere Transcribe supports 14 languages: English, German, French, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Vietnamese, Chinese, Arabic, Japanese, and Korean. It accepts FLAC, MP3, MPEG, MPGA, OGG, and WAV files up to 25 MB.
What is Cohere Transcribe used for?
Cohere Transcribe is a specialized automatic speech recognition model that converts uploaded audio files into text. It is designed for use cases such as meeting transcription, interviews, voice notes, call recordings, and multilingual enterprise speech processing.


Sources 8
Provider

About Cohere