What is Voxtral Mini Transcribe 2?
Voxtral Mini Transcribe 2 is Mistral AI’s batch speech-to-text model. Its primary job is to take an audio recording and return a textual transcript. The model is available as a Premier hosted service through Mistral’s Audio Transcriptions API, with the canonical model identifier voxtral-mini-2602. Mistral’s current transcription documentation also exposes the rolling identifier voxtral-mini-latest; this is an API alias rather than a separate model.
Unlike a general-purpose conversational model, Voxtral Mini Transcribe 2 is focused on recognizing spoken language in supplied recordings. It does not generate images, video, music, or speech, and the supplied specifications do not identify general reasoning, coding, function-calling, or web-search features. Its output is text, with optional speaker and timing information that makes the transcript easier to analyze or synchronize with the original audio.
Where it fits in Mistral’s audio lineup
Mistral positions Voxtral Mini Transcribe 2 for bounded, offline-style transcription jobs. A separate model, Voxtral Mini Transcribe Realtime, is intended for streaming applications and live, low-latency audio. That distinction is important: Voxtral Mini Transcribe 2 is a better fit when an application can submit a recording for processing, while a realtime model is more appropriate when words must be recognized continuously during a conversation or live event.
The model can process recordings of up to three hours per request. This makes it relevant to full meetings, interviews, lectures, calls, and other long-form audio without requiring the recording to be treated as a collection of very short clips. The research does not specify a context-window size or a maximum number of output tokens; the documented practical input limit is the three-hour recording limit.
Supported languages and transcription features
Voxtral Mini Transcribe 2 supports 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch. This makes it useful for multilingual organizations and audio libraries that contain more than English speech.
- Speaker diarization: the model can distinguish and label different speakers, including start and end timing information. This is useful for interviews, meetings, panels, and recorded calls where a plain block of text is not enough.
- Context biasing: an application can provide up to 100 words or phrases to guide recognition of names, product terms, technical vocabulary, or other domain-specific language. Mistral says this feature is optimized for English; support for other languages is experimental.
- Word-level timestamps: individual words can include precise timing, helping applications align transcripts with audio, create subtitles, or make recordings searchable at specific playback positions.
- Noise robustness: Mistral describes the model as designed to maintain transcription quality in challenging acoustic environments. This should be treated as a provider claim rather than a guarantee for every recording.
- Long recordings: a single request can cover up to three hours of audio, according to the supplied model research.
These features make the model more useful than a basic speech-to-text endpoint when the transcript must be attributed to speakers, synchronized with audio, or adapted to specialized terminology. They do not eliminate the need for review: poor audio, unusual accents, background noise, and overlapping voices can still affect the result.
Pricing and API access
Mistral lists standard hosted inference for Voxtral Mini Transcribe 2 at $0.003 per minute. The official pricing information also lists $0.0003 per minute in the batch column. The supplied research does not identify a separate output-token charge.
At the standard rate, one hour of audio would cost approximately $0.18 before any applicable account, billing, or usage conditions. The batch price is lower where the workload qualifies for Mistral’s batch processing option, but the exact operational conditions for that pricing should be confirmed in the current Mistral billing documentation.
The model is available through Mistral’s /v1/audio/transcriptions endpoint. The research confirms the endpoint and model identifiers but does not provide a current SDK example, so implementation details such as upload format, response fields, authentication, and language-specific parameters should be checked against the live API documentation rather than inferred here.
Input, output, and capability profile
| Specification | Verified detail |
|---|---|
| Provider | Mistral AI |
| Model identifier | voxtral-mini-2602 |
| Rolling alias | voxtral-mini-latest |
| Primary input | Audio recordings |
| Primary output | Text transcription |
| Languages | 13 |
| Maximum recording length | Up to three hours per request |
| Speaker diarization | Supported |
| Context biasing | Up to 100 words or phrases |
| Word-level timestamps | Supported |
| Streaming | Not supported for this model |
| Standard price | $0.003 per minute |
| Batch price listing | $0.0003 per minute |
Voxtral Mini Transcribe 2 has audio input and text output. It should not be described as a general multimodal model simply because it handles audio: the documented task is transcription, not open-ended analysis across text, images, video, and audio. The supplied specifications also mark tool use, fine-tuning, JSON mode, caching, and streaming as unavailable or unsupported for this model.
Main strengths and trade-offs
The model’s strongest advantage is task specialization. A transcription-focused service can be easier to evaluate and operate than a general language model being asked to infer transcripts through a broader audio workflow. Voxtral Mini Transcribe 2 combines multilingual recognition with features that matter in production transcript pipelines: speaker labels, custom vocabulary, word timing, and a three-hour request limit.
Its price is another practical strength. The standard rate of $0.003 per minute is low enough for processing substantial archives, while the listed batch rate can reduce cost for workloads that do not require immediate results. The trade-off is that batch processing is inherently less suitable for interactive applications, and the model does not provide the realtime behavior associated with streaming transcription systems.
Diarization is useful but should not be treated as perfect speaker separation. Mistral notes that when voices overlap, the model typically transcribes one speaker. Recordings with frequent interruptions or simultaneous speech may therefore require manual correction or additional processing.
Best use cases
- Meetings and interviews: speaker labels and word timing make it easier to determine who said what and locate statements in the recording.
- Call-center archives: organizations can convert recorded calls into searchable text for review, quality processes, or internal analysis.
- Subtitles and captions: word-level timestamps provide a foundation for synchronizing text with recorded media.
- Compliance and records management: long recordings can be converted into text for document retention and review workflows, subject to the organization’s privacy and governance requirements.
- Multilingual media processing: support for 13 languages helps process international interviews, customer recordings, and media libraries.
- Searchable audio: timestamped transcripts can let users search a recording and jump to the relevant section.
- Specialized terminology: context biasing can improve recognition of names, technical terms, and organization-specific vocabulary, particularly in English.
When to choose Voxtral Mini Transcribe 2
Choose Voxtral Mini Transcribe 2 when the input is an existing recording, immediate word-by-word response is unnecessary, and the resulting transcript benefits from multilingual support, speaker attribution, custom vocabulary, or precise timing. It is especially well matched to scheduled processing and archive-scale transcription where cost per minute matters.
Choose a realtime transcription option instead when the application must display or act on speech while someone is speaking. Mistral specifically distinguishes Voxtral Mini Transcribe Realtime for streaming workloads. A general-purpose language model may also be more appropriate after transcription if the main task is summarization, reasoning, coding, or interactive question answering rather than speech recognition itself.
For heavily overlapping dialogue, plan for transcript review because the model may transcribe only one speaker during overlap. For non-English context biasing, test the results carefully because Mistral identifies that capability as experimental outside English. These are practical evaluation points rather than reasons to reject the model outright.
Limitations to consider
Voxtral Mini Transcribe 2 is not a general conversational or reasoning model. Its documented output is transcription text, and the supplied specifications do not support claims about coding, tool calls, structured JSON output, speech generation, or image and video handling. It also does not provide a documented context length or maximum output-token limit beyond the stated three-hour audio limit.
The model is not designed for low-latency voice agents or sub-200-millisecond streaming. Its batch orientation creates a natural delay between submitting a recording and receiving a transcript. Accuracy can also vary with recording quality, background noise, accents, specialized vocabulary, and overlapping speech, so important transcripts should be reviewed before being used as authoritative records.
Overall, Voxtral Mini Transcribe 2 is best understood as a cost-conscious, feature-rich batch transcription service rather than an all-purpose audio model. Its combination of 13-language support, three-hour recordings, diarization, context biasing, and word-level timestamps gives it a clear role in recorded-audio workflows, while realtime applications should use a streaming-oriented alternative.

