What is Cohere Transcribe Arabic?
Cohere Transcribe Arabic is a specialized automatic speech recognition (ASR) model from Cohere. ASR models convert spoken audio into written text; unlike a general-purpose language model, this model is not intended to hold conversations, generate images, write software, or answer questions from a knowledge base.
The model contains approximately 2 billion parameters and is a fine-tuned version of Cohere Transcribe. Its Conformer-based encoder-decoder architecture is designed to process audio waveforms and produce text. Cohere released it on July 7, 2026, with a focus on Arabic speech, including regional dialects and recordings in which speakers move between Arabic and English.
For users, the important distinction is that this is an Arabic-focused transcription engine rather than a general multilingual assistant. It is useful when the central task is turning recorded speech into searchable, editable text.
Where it fits in Cohere's catalog
Cohere is best known for enterprise-oriented generative AI and retrieval products, including its Command, Embed, Rerank, Parse, Aya, and North offerings. Cohere Transcribe Arabic extends that ecosystem into speech recognition. It is exposed through Cohere's audio transcription API, listed for deployment through Model Vault, and distributed as open weights on Hugging Face.
The Arabic model is a specialized member of the Cohere Transcribe family. Cohere's documentation states that this Arabic variant is not yet available on other cloud platforms. That makes the API, open-weight deployment, and Model Vault the relevant availability paths identified in the supplied documentation.
Languages and audio support
The model supports Arabic and English audio. Its Arabic coverage is intended to include major regional dialects rather than only one standardized variety. It also supports Arabic-English code-switching, such as a conversation in which a speaker changes languages within the same sentence or discussion.
This combination is particularly relevant to customer-support calls, interviews, meetings, broadcasts, and other recordings from multilingual environments. The model is also described as being optimized for difficult far-field conditions, where the microphone is not close to the speaker and the audio may be less clean than a studio recording.
The API accepts FLAC, MP3, MPEG, MPGA, OGG, and WAV files. The documented maximum file size is 25 MB. The supplied research does not specify a maximum duration, context window, token limit, or maximum number of words in the returned transcript, so those values should not be assumed from the file-size limit.
How the model is used
Cohere provides access through its V2 Audio Transcriptions API. The request supplies an audio file and the required language information, and the response returns text. The documented model identifier is cohere-transcribe-arabic-07-2026.
The model can also be used outside Cohere-hosted API infrastructure. Its weights are available under the Apache 2.0 license, which can make self-managed deployment more practical for teams that need control over infrastructure or want to integrate transcription into an existing processing system. Model Vault is another deployment route, with pricing based on hourly instances and availability managed through Cohere.
The open-weight option does not remove operational requirements. Teams running the model themselves remain responsible for suitable hardware, inference software, scaling, audio handling, monitoring, and security. The supplied research does not state a minimum hardware configuration or a guaranteed throughput figure.
Main strengths
- Arabic specialization: The model is designed specifically for Arabic audio instead of treating Arabic as a secondary use case in a broad speech model.
- Dialect coverage: It targets major regional Arabic dialects, which can help with real-world speech that differs from formal written Arabic.
- Code-switching: Arabic-English conversations and utterances are supported, making the model more suitable for multilingual workplaces and customer interactions.
- Far-field focus: Cohere describes the model as optimized for challenging recordings in which the speaker is not directly next to the microphone.
- Deployment flexibility: Users can choose the Cohere API, open weights under Apache 2.0, or Model Vault for an enterprise deployment.
- Production-oriented access: The model is exposed through a dedicated transcription API and is positioned for high-throughput inference rather than conversational interaction.
These strengths are most meaningful when Arabic accuracy, dialect handling, and deployment control matter more than broad audio features or general-purpose reasoning.
Limitations to consider
Cohere Transcribe Arabic returns text only. It does not generate speech, music, images, or video. It is not a general-purpose language model and should not be selected for document drafting, coding, open-ended reasoning, or tool-driven workflows after transcription unless a separate model is added to the pipeline.
The supplied specifications also identify two important transcription limitations: the model does not provide timestamps and does not offer automatic speaker diarization. A transcript therefore does not automatically indicate when each word was spoken or which participant said each passage. Teams building subtitles, searchable media with precise time navigation, meeting minutes separated by speaker, or call analytics may need additional alignment and diarization components.
The 25 MB upload limit may also affect long recordings or high-bitrate files. Large recordings may need to be compressed or divided into segments, subject to the application's quality and continuity requirements. The research does not specify an official maximum recording duration or a built-in method for stitching segmented transcripts together.
Pricing and availability
Cohere offers access through its API for experimentation at no charge, subject to rate limits. The trial access is intended for testing and is not permitted for production or commercial use according to the supplied availability notes. Production access requires moving through Cohere's account and billing process.
Model Vault deployment is priced per hour-instance and requires contacting Cohere. No public numeric price is supplied for either production API use or Model Vault in the research provided here. The open-weight release under the Apache 2.0 license may reduce licensing barriers for self-managed use, but infrastructure and operating costs still apply.
Because there is no verified per-minute, per-file, or per-token price in the supplied material, a direct cost comparison with other transcription services would be speculative. The available editorial assessment rates the model highly for speed and cost relative to its specialized purpose, but those are editorial scores, not provider-published benchmarks or prices.
Capabilities and trade-offs
| Area | What is documented |
|---|---|
| Input | Audio waveforms in FLAC, MP3, MPEG, MPGA, OGG, or WAV format |
| Output | Text transcription |
| Languages | Arabic and English, including Arabic-English code-switching |
| File limit | Up to 25 MB per audio file |
| Architecture | Approximately 2B parameters; Conformer-based encoder-decoder |
| Timestamps | Not provided |
| Speaker diarization | Not provided automatically |
| Tool or function calling | Not supported as a model capability |
| Streaming | Not documented as supported in the supplied research |
| Fine-tuning | Not verified in the supplied research |
Reasoning and coding are not meaningful strengths of this model. The editorial capability fields assign low scores in those areas because they are outside the model's intended purpose, not because the model is being evaluated as a failed coding or reasoning system. Similarly, the model's speed and cost scores are editorial estimates for a specialized ASR workload rather than published Cohere measurements.
Best use cases
Cohere Transcribe Arabic is a strong candidate for applications where Arabic speech must be converted into text at scale:
- Transcribing Arabic customer-support and call-center recordings
- Creating searchable text from Arabic meetings, interviews, lectures, or broadcasts
- Processing conversations that switch between Arabic and English
- Building Arabic-language archives, research collections, or internal knowledge systems
- Running transcription in a controlled environment using open weights or Model Vault
- Handling recordings captured at a distance from the speaker, where far-field performance is important
A common production design could use the model for the first transcription stage, then pass the resulting text to separate software for punctuation refinement, timestamps, speaker labeling, translation, summarization, or search indexing. Those additional functions should not be attributed to Transcribe Arabic itself.
When to choose this model
Choose Cohere Transcribe Arabic when Arabic is central to the workload, dialect and code-switching coverage are important, and you want a choice between hosted API access and controlled deployment. It is especially suitable for organizations that value open weights, Apache 2.0 licensing, or enterprise deployment through Model Vault.
Another speech model may be more appropriate when you need automatic speaker diarization, word- or segment-level timestamps, a documented streaming mode, broader audio event understanding, or a public usage price that can be calculated directly before integration. A general-purpose language model should be added after transcription, or selected instead, when the main task is reasoning over text, generating content, writing code, or calling tools.
The practical trade-off is specialization versus breadth. Cohere Transcribe Arabic concentrates its capabilities on Arabic and English speech recognition, which can be preferable to a broader but less targeted option for Arabic-heavy workloads. In return, it does not provide the surrounding editing, analysis, timing, or speaker-management features that some end-to-end transcription platforms include.
Bottom line
Cohere Transcribe Arabic is a focused 2B-parameter ASR model for Arabic and English audio, with particular attention to dialect variation, Arabic-English code-switching, and far-field recordings. Its combination of API access, open weights, and Model Vault deployment gives organizations several ways to integrate it. The main reasons to look elsewhere are equally clear: no timestamps, no automatic speaker diarization, no general-purpose reasoning or coding, and no publicly supplied production price in the available documentation.

