Cohere Transcribe

Cohere Transcribe Arabic

by Cohere · Live

A specialized 2B-parameter Conformer-based speech recognition model for Arabic and English audio. It supports regional Arabic dialects, Arabic-English code-switching and challenging far-field recordings, accepts files up to 25 MB, and returns text through Cohere's transcription API. Open weights under Apache 2.0 and Model Vault deployment are also available, but the model does not provide timestamps or automatic speaker diarization.

Text Reasoning Coding
Cohere Transcribe Arabic is an audio-in, text-out speech recognition model released on July 7, 2026. It is designed for Arabic transcription rather than general-purpose text generation, with support for Arabic dialect variation, Arabic-English code-switching, and production-oriented audio workloads. The model can be used through Cohere's V2 Audio Transcriptions API or deployed from its open weights, while Model Vault provides an enterprise deployment option.
Outputs

What Cohere Transcribe Arabic can produce

Text
Inputs

What it can understand

Audio
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Cohere Transcribe
Model type Other
Release date 2026-07-07
Status Live
Knowledge cutoff notes

A conventional textual knowledge cutoff is not applicable to this dedicated automatic speech recognition model. The model transcribes supplied audio rather than answering from a documented pretraining knowledge date.

Model notes

Exact model ID: cohere-transcribe-arabic-07-2026. This is a fine-tuned version of Cohere Transcribe and contains approximately 2B parameters. It uses a Conformer-based encoder-decoder architecture, accepts audio waveform input and returns text. Supported languages are Arabic and English, with multidialectal and Arabic-English code-switching support. The API accepts FLAC, MP3, MPEG, MPGA, OGG and WAV files up to 25MB. The model does not provide timestamps or automatic speaker diarization. It is available through Cohere’s V2 Audio Transcriptions API, as open weights on Hugging Face under the Apache 2.0 license, and through Model Vault. Cohere states that the Arabic variant is not yet available on other cloud platforms. Speed and cost scores are editorial estimates for a specialized ASR model, while reasoning and coding scores reflect that those capabilities are outside the model’s intended purpose.

Cost

Model pricing

Input Free through Cohere API for experimentation subject to rate limits; Model Vault deployment is priced per hour-instance and requires contacting Cohere.
Model guide

Cohere Transcribe Arabic: Open-Weight ASR for Arabic Dialects and Code-Switched Audio

Cohere Transcribe Arabic is a 2-billion-parameter automatic speech recognition model specialized for Arabic audio. It transcribes Arabic and English, supports major Arabic dialects and Arabic-English code-switching, handles challenging far-field recordings, and is available through Cohere's Audio Transcriptions API, as open weights under the Apache 2.0 license, and through Model Vault.

What is Cohere Transcribe Arabic?

Cohere Transcribe Arabic is a specialized automatic speech recognition (ASR) model from Cohere. ASR models convert spoken audio into written text; unlike a general-purpose language model, this model is not intended to hold conversations, generate images, write software, or answer questions from a knowledge base.

The model contains approximately 2 billion parameters and is a fine-tuned version of Cohere Transcribe. Its Conformer-based encoder-decoder architecture is designed to process audio waveforms and produce text. Cohere released it on July 7, 2026, with a focus on Arabic speech, including regional dialects and recordings in which speakers move between Arabic and English.

For users, the important distinction is that this is an Arabic-focused transcription engine rather than a general multilingual assistant. It is useful when the central task is turning recorded speech into searchable, editable text.

Where it fits in Cohere's catalog

Cohere is best known for enterprise-oriented generative AI and retrieval products, including its Command, Embed, Rerank, Parse, Aya, and North offerings. Cohere Transcribe Arabic extends that ecosystem into speech recognition. It is exposed through Cohere's audio transcription API, listed for deployment through Model Vault, and distributed as open weights on Hugging Face.

The Arabic model is a specialized member of the Cohere Transcribe family. Cohere's documentation states that this Arabic variant is not yet available on other cloud platforms. That makes the API, open-weight deployment, and Model Vault the relevant availability paths identified in the supplied documentation.

Languages and audio support

The model supports Arabic and English audio. Its Arabic coverage is intended to include major regional dialects rather than only one standardized variety. It also supports Arabic-English code-switching, such as a conversation in which a speaker changes languages within the same sentence or discussion.

This combination is particularly relevant to customer-support calls, interviews, meetings, broadcasts, and other recordings from multilingual environments. The model is also described as being optimized for difficult far-field conditions, where the microphone is not close to the speaker and the audio may be less clean than a studio recording.

The API accepts FLAC, MP3, MPEG, MPGA, OGG, and WAV files. The documented maximum file size is 25 MB. The supplied research does not specify a maximum duration, context window, token limit, or maximum number of words in the returned transcript, so those values should not be assumed from the file-size limit.

How the model is used

Cohere provides access through its V2 Audio Transcriptions API. The request supplies an audio file and the required language information, and the response returns text. The documented model identifier is cohere-transcribe-arabic-07-2026.

The model can also be used outside Cohere-hosted API infrastructure. Its weights are available under the Apache 2.0 license, which can make self-managed deployment more practical for teams that need control over infrastructure or want to integrate transcription into an existing processing system. Model Vault is another deployment route, with pricing based on hourly instances and availability managed through Cohere.

The open-weight option does not remove operational requirements. Teams running the model themselves remain responsible for suitable hardware, inference software, scaling, audio handling, monitoring, and security. The supplied research does not state a minimum hardware configuration or a guaranteed throughput figure.

Main strengths

  • Arabic specialization: The model is designed specifically for Arabic audio instead of treating Arabic as a secondary use case in a broad speech model.
  • Dialect coverage: It targets major regional Arabic dialects, which can help with real-world speech that differs from formal written Arabic.
  • Code-switching: Arabic-English conversations and utterances are supported, making the model more suitable for multilingual workplaces and customer interactions.
  • Far-field focus: Cohere describes the model as optimized for challenging recordings in which the speaker is not directly next to the microphone.
  • Deployment flexibility: Users can choose the Cohere API, open weights under Apache 2.0, or Model Vault for an enterprise deployment.
  • Production-oriented access: The model is exposed through a dedicated transcription API and is positioned for high-throughput inference rather than conversational interaction.

These strengths are most meaningful when Arabic accuracy, dialect handling, and deployment control matter more than broad audio features or general-purpose reasoning.

Limitations to consider

Cohere Transcribe Arabic returns text only. It does not generate speech, music, images, or video. It is not a general-purpose language model and should not be selected for document drafting, coding, open-ended reasoning, or tool-driven workflows after transcription unless a separate model is added to the pipeline.

The supplied specifications also identify two important transcription limitations: the model does not provide timestamps and does not offer automatic speaker diarization. A transcript therefore does not automatically indicate when each word was spoken or which participant said each passage. Teams building subtitles, searchable media with precise time navigation, meeting minutes separated by speaker, or call analytics may need additional alignment and diarization components.

The 25 MB upload limit may also affect long recordings or high-bitrate files. Large recordings may need to be compressed or divided into segments, subject to the application's quality and continuity requirements. The research does not specify an official maximum recording duration or a built-in method for stitching segmented transcripts together.

Pricing and availability

Cohere offers access through its API for experimentation at no charge, subject to rate limits. The trial access is intended for testing and is not permitted for production or commercial use according to the supplied availability notes. Production access requires moving through Cohere's account and billing process.

Model Vault deployment is priced per hour-instance and requires contacting Cohere. No public numeric price is supplied for either production API use or Model Vault in the research provided here. The open-weight release under the Apache 2.0 license may reduce licensing barriers for self-managed use, but infrastructure and operating costs still apply.

Because there is no verified per-minute, per-file, or per-token price in the supplied material, a direct cost comparison with other transcription services would be speculative. The available editorial assessment rates the model highly for speed and cost relative to its specialized purpose, but those are editorial scores, not provider-published benchmarks or prices.

Capabilities and trade-offs

AreaWhat is documented
InputAudio waveforms in FLAC, MP3, MPEG, MPGA, OGG, or WAV format
OutputText transcription
LanguagesArabic and English, including Arabic-English code-switching
File limitUp to 25 MB per audio file
ArchitectureApproximately 2B parameters; Conformer-based encoder-decoder
TimestampsNot provided
Speaker diarizationNot provided automatically
Tool or function callingNot supported as a model capability
StreamingNot documented as supported in the supplied research
Fine-tuningNot verified in the supplied research

Reasoning and coding are not meaningful strengths of this model. The editorial capability fields assign low scores in those areas because they are outside the model's intended purpose, not because the model is being evaluated as a failed coding or reasoning system. Similarly, the model's speed and cost scores are editorial estimates for a specialized ASR workload rather than published Cohere measurements.

Best use cases

Cohere Transcribe Arabic is a strong candidate for applications where Arabic speech must be converted into text at scale:

  • Transcribing Arabic customer-support and call-center recordings
  • Creating searchable text from Arabic meetings, interviews, lectures, or broadcasts
  • Processing conversations that switch between Arabic and English
  • Building Arabic-language archives, research collections, or internal knowledge systems
  • Running transcription in a controlled environment using open weights or Model Vault
  • Handling recordings captured at a distance from the speaker, where far-field performance is important

A common production design could use the model for the first transcription stage, then pass the resulting text to separate software for punctuation refinement, timestamps, speaker labeling, translation, summarization, or search indexing. Those additional functions should not be attributed to Transcribe Arabic itself.

When to choose this model

Choose Cohere Transcribe Arabic when Arabic is central to the workload, dialect and code-switching coverage are important, and you want a choice between hosted API access and controlled deployment. It is especially suitable for organizations that value open weights, Apache 2.0 licensing, or enterprise deployment through Model Vault.

Another speech model may be more appropriate when you need automatic speaker diarization, word- or segment-level timestamps, a documented streaming mode, broader audio event understanding, or a public usage price that can be calculated directly before integration. A general-purpose language model should be added after transcription, or selected instead, when the main task is reasoning over text, generating content, writing code, or calling tools.

The practical trade-off is specialization versus breadth. Cohere Transcribe Arabic concentrates its capabilities on Arabic and English speech recognition, which can be preferable to a broader but less targeted option for Arabic-heavy workloads. In return, it does not provide the surrounding editing, analysis, timing, or speaker-management features that some end-to-end transcription platforms include.

Bottom line

Cohere Transcribe Arabic is a focused 2B-parameter ASR model for Arabic and English audio, with particular attention to dialect variation, Arabic-English code-switching, and far-field recordings. Its combination of API access, open weights, and Model Vault deployment gives organizations several ways to integrate it. The main reasons to look elsewhere are equally clear: no timestamps, no automatic speaker diarization, no general-purpose reasoning or coding, and no publicly supplied production price in the available documentation.


Answers to Frequently Asked Questions

Is Cohere Transcribe Arabic free to use?
Cohere provides no-charge API access for experimentation subject to rate limits, but the supplied availability information says this trial access is not permitted for production or commercial use. Production API pricing and Model Vault pricing are not publicly specified in the available documentation, while self-hosting still incurs infrastructure and operating costs.
How can Cohere Transcribe Arabic be deployed?
The model is available through Cohere's V2 Audio Transcriptions API, through Model Vault, and as open weights on Hugging Face under the Apache 2.0 license. Self-managed deployment provides infrastructure control but requires teams to handle hardware, inference software, scaling, monitoring, audio processing, and security.
Does Cohere Transcribe Arabic provide timestamps or speaker diarization?
No. Cohere Transcribe Arabic does not automatically provide timestamps or speaker diarization. Applications that need word or segment timing, speaker labels, subtitles, or speaker-separated meeting notes must add separate alignment and diarization components.
What is Cohere Transcribe Arabic used for?
Cohere Transcribe Arabic is an automatic speech recognition model designed to convert Arabic and English audio into text. It is suited to Arabic customer-support calls, meetings, interviews, lectures, broadcasts, archives, and other recordings, including audio with Arabic-English code-switching.
Which Arabic dialects and audio formats does Cohere Transcribe Arabic support?
The model is designed to support major regional Arabic dialects as well as English and Arabic-English code-switching. Its API accepts FLAC, MP3, MPEG, MPGA, OGG, and WAV files, with a documented maximum file size of 25 MB.


Sources 7
Provider

About Cohere