Voxtral

Voxtral Mini Transcribe 2

by Mistral AI · Active; GA; Premier hosted model

Mistral AI’s Voxtral Mini Transcribe 2 converts recordings into text through a hosted batch transcription API. It supports 13 languages, audio up to three hours per request, speaker diarization, custom vocabulary guidance, and word-level timestamps. Its standard price is $0.003 per minute, with a lower batch price listed by Mistral. The model is intended for recorded audio rather than realtime voice applications.

Text Reasoning Coding
Voxtral Mini Transcribe 2 is a specialized audio-input model from Mistral AI for offline and batch speech transcription. It converts recorded speech into text, can identify different speakers, accepts custom vocabulary guidance, and provides timing information down to individual words. The model is aimed at processing completed recordings rather than powering a live voice conversation or low-latency audio agent.
Outputs

What Voxtral Mini Transcribe 2 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Voxtral
Model type Other
Release date 2026-02-04
Status Active; GA; Premier hosted model
Knowledge cutoff notes

Mistral does not publish a model-specific knowledge cutoff for Voxtral Mini Transcribe 2. As a transcription model, its primary behavior is based on supplied audio rather than a documented general-purpose textual knowledge base.

Model notes

Canonical model identifier is voxtral-mini-2602. Current transcription documentation also exposes the model through voxtral-mini-latest, which is a rolling API alias rather than a separate model. Supports 13 languages, speaker diarization, context biasing with up to 100 words or phrases, word-level timestamps, and recordings up to 3 hours per request. Context biasing is optimized for English and experimental for other languages. Mistral notes that overlapping speech is typically handled by transcribing one speaker. Official pricing lists $0.003 per minute for standard inference and $0.0003 per minute in the batch column; no separate output price is listed. It is a transcription-only service and should not be confused with Voxtral Mini Transcribe Realtime.

Cost

Model pricing

Input $0.003 per minute
Model guide

Voxtral Mini Transcribe 2: Mistral’s Batch Speech-to-Text Model for Long Recordings

Voxtral Mini Transcribe 2 is Mistral AI’s hosted batch speech-recognition model for converting recordings into text. It supports 13 languages, recordings up to three hours per request, speaker diarization, context biasing, and word-level timestamps. Its $0.003-per-minute standard inference price makes it suitable for meetings, interviews, calls, subtitles, compliance archives, and searchable audio, but it is not designed for realtime voice applications.

What is Voxtral Mini Transcribe 2?

Voxtral Mini Transcribe 2 is Mistral AI’s batch speech-to-text model. Its primary job is to take an audio recording and return a textual transcript. The model is available as a Premier hosted service through Mistral’s Audio Transcriptions API, with the canonical model identifier voxtral-mini-2602. Mistral’s current transcription documentation also exposes the rolling identifier voxtral-mini-latest; this is an API alias rather than a separate model.

Unlike a general-purpose conversational model, Voxtral Mini Transcribe 2 is focused on recognizing spoken language in supplied recordings. It does not generate images, video, music, or speech, and the supplied specifications do not identify general reasoning, coding, function-calling, or web-search features. Its output is text, with optional speaker and timing information that makes the transcript easier to analyze or synchronize with the original audio.

Where it fits in Mistral’s audio lineup

Mistral positions Voxtral Mini Transcribe 2 for bounded, offline-style transcription jobs. A separate model, Voxtral Mini Transcribe Realtime, is intended for streaming applications and live, low-latency audio. That distinction is important: Voxtral Mini Transcribe 2 is a better fit when an application can submit a recording for processing, while a realtime model is more appropriate when words must be recognized continuously during a conversation or live event.

The model can process recordings of up to three hours per request. This makes it relevant to full meetings, interviews, lectures, calls, and other long-form audio without requiring the recording to be treated as a collection of very short clips. The research does not specify a context-window size or a maximum number of output tokens; the documented practical input limit is the three-hour recording limit.

Supported languages and transcription features

Voxtral Mini Transcribe 2 supports 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch. This makes it useful for multilingual organizations and audio libraries that contain more than English speech.

  • Speaker diarization: the model can distinguish and label different speakers, including start and end timing information. This is useful for interviews, meetings, panels, and recorded calls where a plain block of text is not enough.
  • Context biasing: an application can provide up to 100 words or phrases to guide recognition of names, product terms, technical vocabulary, or other domain-specific language. Mistral says this feature is optimized for English; support for other languages is experimental.
  • Word-level timestamps: individual words can include precise timing, helping applications align transcripts with audio, create subtitles, or make recordings searchable at specific playback positions.
  • Noise robustness: Mistral describes the model as designed to maintain transcription quality in challenging acoustic environments. This should be treated as a provider claim rather than a guarantee for every recording.
  • Long recordings: a single request can cover up to three hours of audio, according to the supplied model research.

These features make the model more useful than a basic speech-to-text endpoint when the transcript must be attributed to speakers, synchronized with audio, or adapted to specialized terminology. They do not eliminate the need for review: poor audio, unusual accents, background noise, and overlapping voices can still affect the result.

Pricing and API access

Mistral lists standard hosted inference for Voxtral Mini Transcribe 2 at $0.003 per minute. The official pricing information also lists $0.0003 per minute in the batch column. The supplied research does not identify a separate output-token charge.

At the standard rate, one hour of audio would cost approximately $0.18 before any applicable account, billing, or usage conditions. The batch price is lower where the workload qualifies for Mistral’s batch processing option, but the exact operational conditions for that pricing should be confirmed in the current Mistral billing documentation.

The model is available through Mistral’s /v1/audio/transcriptions endpoint. The research confirms the endpoint and model identifiers but does not provide a current SDK example, so implementation details such as upload format, response fields, authentication, and language-specific parameters should be checked against the live API documentation rather than inferred here.

Input, output, and capability profile

SpecificationVerified detail
ProviderMistral AI
Model identifiervoxtral-mini-2602
Rolling aliasvoxtral-mini-latest
Primary inputAudio recordings
Primary outputText transcription
Languages13
Maximum recording lengthUp to three hours per request
Speaker diarizationSupported
Context biasingUp to 100 words or phrases
Word-level timestampsSupported
StreamingNot supported for this model
Standard price$0.003 per minute
Batch price listing$0.0003 per minute

Voxtral Mini Transcribe 2 has audio input and text output. It should not be described as a general multimodal model simply because it handles audio: the documented task is transcription, not open-ended analysis across text, images, video, and audio. The supplied specifications also mark tool use, fine-tuning, JSON mode, caching, and streaming as unavailable or unsupported for this model.

Main strengths and trade-offs

The model’s strongest advantage is task specialization. A transcription-focused service can be easier to evaluate and operate than a general language model being asked to infer transcripts through a broader audio workflow. Voxtral Mini Transcribe 2 combines multilingual recognition with features that matter in production transcript pipelines: speaker labels, custom vocabulary, word timing, and a three-hour request limit.

Its price is another practical strength. The standard rate of $0.003 per minute is low enough for processing substantial archives, while the listed batch rate can reduce cost for workloads that do not require immediate results. The trade-off is that batch processing is inherently less suitable for interactive applications, and the model does not provide the realtime behavior associated with streaming transcription systems.

Diarization is useful but should not be treated as perfect speaker separation. Mistral notes that when voices overlap, the model typically transcribes one speaker. Recordings with frequent interruptions or simultaneous speech may therefore require manual correction or additional processing.

Best use cases

  • Meetings and interviews: speaker labels and word timing make it easier to determine who said what and locate statements in the recording.
  • Call-center archives: organizations can convert recorded calls into searchable text for review, quality processes, or internal analysis.
  • Subtitles and captions: word-level timestamps provide a foundation for synchronizing text with recorded media.
  • Compliance and records management: long recordings can be converted into text for document retention and review workflows, subject to the organization’s privacy and governance requirements.
  • Multilingual media processing: support for 13 languages helps process international interviews, customer recordings, and media libraries.
  • Searchable audio: timestamped transcripts can let users search a recording and jump to the relevant section.
  • Specialized terminology: context biasing can improve recognition of names, technical terms, and organization-specific vocabulary, particularly in English.

When to choose Voxtral Mini Transcribe 2

Choose Voxtral Mini Transcribe 2 when the input is an existing recording, immediate word-by-word response is unnecessary, and the resulting transcript benefits from multilingual support, speaker attribution, custom vocabulary, or precise timing. It is especially well matched to scheduled processing and archive-scale transcription where cost per minute matters.

Choose a realtime transcription option instead when the application must display or act on speech while someone is speaking. Mistral specifically distinguishes Voxtral Mini Transcribe Realtime for streaming workloads. A general-purpose language model may also be more appropriate after transcription if the main task is summarization, reasoning, coding, or interactive question answering rather than speech recognition itself.

For heavily overlapping dialogue, plan for transcript review because the model may transcribe only one speaker during overlap. For non-English context biasing, test the results carefully because Mistral identifies that capability as experimental outside English. These are practical evaluation points rather than reasons to reject the model outright.

Limitations to consider

Voxtral Mini Transcribe 2 is not a general conversational or reasoning model. Its documented output is transcription text, and the supplied specifications do not support claims about coding, tool calls, structured JSON output, speech generation, or image and video handling. It also does not provide a documented context length or maximum output-token limit beyond the stated three-hour audio limit.

The model is not designed for low-latency voice agents or sub-200-millisecond streaming. Its batch orientation creates a natural delay between submitting a recording and receiving a transcript. Accuracy can also vary with recording quality, background noise, accents, specialized vocabulary, and overlapping speech, so important transcripts should be reviewed before being used as authoritative records.

Overall, Voxtral Mini Transcribe 2 is best understood as a cost-conscious, feature-rich batch transcription service rather than an all-purpose audio model. Its combination of 13-language support, three-hour recordings, diarization, context biasing, and word-level timestamps gives it a clear role in recorded-audio workflows, while realtime applications should use a streaming-oriented alternative.


Answers to Frequently Asked Questions

Is Voxtral Mini Transcribe 2 suitable for real-time transcription?
No. Voxtral Mini Transcribe 2 is designed for batch transcription of existing recordings and does not support streaming. For live conversations, voice agents, or applications that need words recognized while someone is speaking, Mistral’s Voxtral Mini Transcribe Realtime model is a more appropriate option.
How much does Voxtral Mini Transcribe 2 cost?
Mistral lists standard hosted inference at $0.003 per minute, which is approximately $0.18 for one hour of audio. A batch price of $0.0003 per minute is also listed for qualifying batch workloads, but the exact conditions should be confirmed in Mistral’s current billing documentation.
What languages and transcription features does Voxtral Mini Transcribe 2 support?
Voxtral Mini Transcribe 2 supports 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch. It also supports speaker diarization, word-level timestamps, and context biasing for up to 100 words or phrases.
What is Voxtral Mini Transcribe 2 used for?
Voxtral Mini Transcribe 2 is Mistral AI’s batch speech-to-text model for converting audio recordings into text. It is designed for meetings, interviews, lectures, calls, subtitles, searchable audio archives, and other offline transcription workflows.
How long of an audio recording can Voxtral Mini Transcribe 2 transcribe?
The model can process recordings of up to three hours per request. It is intended for long-form audio submitted for batch processing rather than continuous, low-latency streaming.


Sources 7
Provider

About Mistral AI