What is Cohere Transcribe?
Cohere Transcribe is Cohere's automatic speech recognition (ASR) model. In practical terms, it takes an audio file containing spoken language and returns the words as text. The canonical model identifier is cohere-transcribe-03-2026.
The model is an open-source research release with approximately 2 billion parameters and is distributed under the Apache 2.0 license. It is available through Cohere's Audio Transcriptions API for hosted experimentation and application integration, and through Cohere Model Vault for production deployments that require dedicated infrastructure. This places it in Cohere's catalog as a specialized audio model alongside, rather than as a replacement for, the company's text-generation, retrieval, embedding, and document-processing systems.
Unlike a general multimodal language model, Cohere Transcribe has a narrow purpose: converting speech into text. Its documented output is text, and it does not generate audio, images, video, or general-purpose written responses.
Supported languages and audio input
Cohere Transcribe supports 14 languages:
- English
- German
- French
- Italian
- Spanish
- Portuguese
- Greek
- Dutch
- Polish
- Vietnamese
- Chinese
- Arabic
- Japanese
- Korean
The transcription request requires the caller to provide the input language using an ISO-639-1 code. The model is optimized for one specified language at a time, and the supplied documentation does not describe explicit automatic language detection. Applications handling recordings in unknown or mixed languages therefore need to identify the language before transcription or add a separate language-identification step.
The documented API accepts FLAC, MP3, MPEG, MPGA, OGG, and WAV files. The maximum documented upload size is 25 MB. This makes the model suitable for ordinary meeting recordings, interviews, voice notes, and call segments, but larger recordings may need to be divided before submission. The supplied documentation describes an uploaded-file workflow rather than a live streaming transcription interface.
Architecture and performance positioning
Cohere describes Transcribe as using a speech-optimized Conformer-based encoder-decoder architecture. A Conformer combines techniques suited to speech processing with transformer-style sequence modeling. At a high level, the system converts the audio waveform into log-Mel spectrogram features, processes those features with a Conformer encoder, and uses a lightweight Transformer decoder to generate text tokens.
Cohere says the model was trained from scratch with supervised cross-entropy and optimized for low word error rate and efficient serving. These are provider descriptions of the model's design and goals, not an independent benchmark result. Cohere also states that Transcribe can achieve a real-time factor up to three times faster than other dedicated ASR models in the same size range. The claim indicates a focus on throughput and latency, but actual performance will depend on audio characteristics, hardware, deployment configuration, and workload.
The model's relatively focused architecture is important to its positioning. It is not intended to reason over a transcript, write a report, call tools, or answer questions about the recording in the same request. A workflow can transcribe audio first and then send the resulting text to another model or application for summarization, search, classification, or extraction.
API and deployment options
The Audio Transcriptions API provides the simplest documented way to test or integrate Cohere Transcribe. A caller uploads a supported audio file, specifies the language, and receives the transcribed text. Cohere provides free, low-setup API experimentation subject to rate limits. The research does not provide a per-minute, per-input-token, or per-output-token price for the hosted API.
For production use without the trial rate limits, Cohere makes the model available through Model Vault. Model Vault pricing is calculated per hour-instance rather than by input or output token. Cohere also offers longer-term commitment discounts, but the supplied research does not include a public numeric hourly price. Organizations should therefore request or confirm a deployment quote instead of estimating cost from token-based language-model pricing.
Model Vault is relevant when an organization needs a more controlled serving arrangement or predictable dedicated capacity. The hosted API is more appropriate for initial evaluation and applications that can work within its file and rate limits. Neither option is documented here as providing live streaming transcription, so real-time voice applications should verify current product support or consider a service specifically designed for streaming audio.
Capabilities and limitations
The following distinctions summarize what is documented for the model:
| Area | Documented behavior |
|---|---|
| Primary task | Automatic speech recognition and audio-to-text transcription |
| Input | Audio waveform files |
| Output | Text |
| Languages | 14 specified languages |
| File formats | FLAC, MP3, MPEG, MPGA, OGG, and WAV |
| Maximum file size | 25 MB |
| Language selection | Caller must specify the language |
| Streaming | No documented live streaming interface |
| Timestamps | Not provided in the supplied specification |
| Speaker diarization | Not provided in the supplied specification |
| License | Apache 2.0 |
There is no documented context window or maximum output-token limit for Cohere Transcribe. The relevant practical limit is the 25 MB maximum upload size. The response is a transcript, so output length will depend on the duration and speech content of the submitted recording rather than on a published general-purpose language-model token limit.
The model has no documented tool or function-calling support, JSON-mode support, caching, batch API, or fine-tuning capability in the supplied research. It also does not provide reasoning or coding capabilities in the usual language-model sense. Those omissions are not necessarily defects for transcription, but they matter when comparing it with a general-purpose model that can process audio and then perform additional tasks directly.
Main strengths and trade-offs
Where Cohere Transcribe is strong
- Focused speech recognition: The model is purpose-built for converting speech into text rather than sharing capacity with unrelated generation tasks.
- Multilingual coverage: Fourteen supported languages cover many common enterprise and international transcription scenarios.
- Efficient deployment: Cohere positions its Conformer-based design for low word error rate and fast inference, including a claimed real-time factor advantage over comparable dedicated ASR models.
- Open licensing: The Apache 2.0 license can be useful for organizations evaluating open-weight deployment and integration options, subject to their own legal and operational review.
- Enterprise deployment path: Model Vault provides a production route based on dedicated hourly instances rather than token billing.
Where it is limited
- No automatic language detection: The application must provide the language code, which complicates mixed-language or unknown-language recordings.
- No documented timestamps: Users needing word-level or segment-level timing may need an additional processing system.
- No speaker diarization: The model does not identify which participant spoke each portion of a recording.
- File-based workflow: The documented API accepts uploaded files, not a live audio stream.
- 25 MB upload ceiling: Long recordings may need preprocessing or segmentation.
- Limited task scope: It transcribes audio but does not summarize, reason over, search, or transform the transcript by itself.
Best use cases
Cohere Transcribe is a practical fit when the central requirement is multilingual transcription of uploaded recordings. Examples include meeting notes, interview archives, voice-note conversion, customer-service or call-center recordings, multilingual enterprise speech repositories, and internal applications that need a text representation before applying search or language analysis.
It is especially suitable when an organization values an open-source release, wants to evaluate a relatively efficient dedicated ASR model, or plans to use Cohere's enterprise deployment options. A typical pipeline might upload a recording, specify its known language, store the returned transcript, and then pass that text to a separate search, summarization, or document workflow.
When to choose Cohere Transcribe
Choose Cohere Transcribe when you need a dedicated speech-to-text component with support for the listed 14 languages, common uploaded audio formats, and a path from rate-limited API experimentation to Model Vault deployment. Its combination of focused architecture, Apache 2.0 licensing, and Cohere's enterprise infrastructure options may be more relevant than a general-purpose multimodal model when transcription throughput and deployment control are the priorities.
Another speech service may be more appropriate if the application requires automatic language identification, live streaming, timestamps, or speaker separation. A general-purpose language model may be a better second-stage choice when the main task is to analyze the content, produce a structured report, answer questions, or generate code after transcription. Cohere Transcribe can serve as the audio-to-text stage in that workflow, but the supplied specifications do not indicate that it performs those later tasks itself.
Bottom line
Cohere Transcribe is a specialized, open-source multilingual ASR model rather than a broad conversational AI system. Its documented strengths are 14-language coverage, support for common audio files up to 25 MB, an efficient Conformer-based design, and access through both Cohere's transcription API and Model Vault. Its main constraints are equally clear: the language must be specified, the workflow is file-based, and timestamps, diarization, automatic language detection, and general reasoning are not documented. For enterprise audio-to-text workloads that fit those boundaries, it offers a focused alternative to using a larger general-purpose model for transcription.

