What is Voxtral Mini Transcribe Realtime?
Voxtral Mini Transcribe Realtime is a dedicated speech-to-text model from Mistral AI. Its job is narrowly defined: receive live audio and produce a text transcription with low latency. Unlike a general conversational model, it is not designed to answer questions, write code, analyze images, or generate spoken audio.
The model's canonical identifier is voxtral-mini-transcribe-realtime-2602. Mistral released it for general availability on February 4, 2026, under the Apache 2.0 license. It is listed in Mistral's audio documentation as the realtime transcription option, alongside separate models for recorded-audio transcription and speech generation.
How the realtime transcription workflow works
Applications connect to Mistral's realtime transcription service through a WebSocket connection. WebSockets keep a persistent two-way connection open, which is useful when audio must be sent continuously rather than uploaded as one completed file.
As audio arrives, the service can emit different events, including session events, text-delta events, completion events, and error events. A text-delta event contains an incremental piece of the recognized transcript. An application can use these partial results to update a caption window, populate a notes panel, or pass spoken text to another component without waiting for the entire recording to finish.
Mistral's examples use signed 16-bit little-endian PCM audio sampled at 16 kHz for microphone input. The supplied research does not establish that this is the model's only supported audio format, so it is best treated as the documented example configuration rather than a universal format requirement.
Main capabilities
- Live audio transcription: converts an incoming audio stream into text while recording is in progress.
- Incremental results: returns transcription updates through streaming events rather than waiting for a complete response.
- WebSocket access: supports a persistent connection suitable for microphone and other continuous-audio workflows.
- Configurable streaming delay: lets developers tune how long the system waits for additional audio context before producing transcription.
- Text-only output: the result is recognized text, not synthesized speech or another media type.
The configurable delay creates a practical responsiveness-versus-context trade-off. A shorter delay can make captions and voice interfaces feel faster, while a longer target delay may give the recognizer more surrounding audio before it emits text. The research does not provide a fixed latency guarantee or a recommended delay for every use case.
Where it fits in Mistral's audio lineup
Voxtral Mini Transcribe Realtime is the live-streaming member of Mistral's documented audio offerings. It should be distinguished from Voxtral Mini Transcribe 2, which Mistral recommends for recorded-file transcription workflows involving features such as speaker diarization, word-level timestamps, and context biasing. It is also distinct from Voxtral TTS, which is intended for turning text into speech.
This positioning matters when selecting a model. Voxtral Mini Transcribe Realtime is appropriate when text needs to appear during an ongoing conversation or recording. A batch transcription model is more suitable when the complete audio file is already available and the application needs post-processing features that the realtime workflow does not support.
Strengths and practical use cases
The model's main strength is specialization. It exposes a focused path from streaming audio to streaming text, without requiring an application to wait for a recording to finish. That makes it a natural fit for:
- Live captions: display spoken content during meetings, presentations, broadcasts, or events.
- Voice interfaces: convert a person's speech into text that a separate language model or application can process.
- Realtime note-taking: build a running transcript while a discussion is taking place.
- Speech-to-speech pipelines: use transcription as the input stage, then send the text to a reasoning model and a text-to-speech system.
- Interactive audio applications: respond to spoken input without waiting for a full audio upload.
For a voice assistant, Voxtral Mini Transcribe Realtime is one component rather than a complete assistant. The application still needs separate logic or a language model to interpret the transcript and decide what to do, plus a text-to-speech service if it must answer aloud.
Limitations and unsupported features
The model is optimized for transcription and should not be evaluated as a general-purpose AI assistant. The supplied specifications do not document general chat, reasoning, coding, image understanding, image generation, audio generation, tool calling, or structured-output capabilities.
Realtime transcription is currently incompatible with the diarize parameter. Diarization identifies which speaker said which words, so applications that require speaker attribution should use Mistral's appropriate non-realtime transcription workflow instead. Voxtral Mini Transcribe Realtime is also not documented as a batch transcription model for completed recordings.
There is no verified context-window size or maximum output-token limit in the supplied documentation. Those limits are less directly applicable to a streaming recognizer than they are to a text-generation model, but developers should still confirm audio-session and service limits before deploying at scale. The research also does not provide benchmark accuracy, guaranteed latency, supported-language coverage, or a maximum concurrent-connection limit.
Inputs, outputs, and technical profile
| Specification | Verified detail |
|---|---|
| Model ID | voxtral-mini-transcribe-realtime-2602 |
| Provider | Mistral AI |
| Input | Streaming audio |
| Documented example audio | PCM signed 16-bit little-endian at 16 kHz |
| Output | Incremental transcription text |
| Connection type | WebSocket realtime endpoint |
| Streaming | Supported |
| Tool or function use | Not documented for this model |
| Fine-tuning | Not documented |
| Context length | Not published in the supplied research |
| Maximum output tokens | Not published in the supplied research |
| License | Apache 2.0 |
The table separates verified specifications from missing information. In particular, the model's streaming behavior is documented, but no claim should be made about a guaranteed response time or an unlimited session length.
Pricing and cost positioning
Mistral lists Voxtral Mini Transcribe Realtime at $0.006 per minute. The supplied research describes this as the audio-minute price, with transcription output included rather than billed as a separate output-token charge. That makes the pricing model straightforward for applications whose main variable is the amount of audio processed.
The cost should be considered alongside the model's specialized scope. A realtime speech recognizer can be more appropriate than sending raw audio to a larger multimodal or conversational system when the required result is simply a transcript. Conversely, an application that needs reasoning, extraction, tool calls, or a spoken response will incur additional costs and complexity because those functions require other components.
Reasoning, coding, and tool support
Voxtral Mini Transcribe Realtime is not a reasoning or coding model. It recognizes speech; it does not independently interpret a user's request, write software, browse the web, call external tools, or produce a final business decision. A developer can place it at the front of a larger pipeline, but those downstream capabilities come from other services or application code.
This distinction prevents a common implementation mistake: treating a transcription event as an assistant response. The text emitted by the model is an intermediate representation of what was heard. Applications should validate partial transcripts and decide whether to wait for a completion event before triggering an irreversible action.
When to choose this model
Choose Voxtral Mini Transcribe Realtime when the defining requirement is text from live audio with incremental updates. It is a strong candidate for captions, live meeting notes, voice-command input, and the transcription stage of a speech-to-speech system. Its WebSocket interface and configurable streaming delay are particularly relevant when an application must balance immediacy against waiting for more audio context.
Choose another option when the audio is already recorded and you need speaker labels, word-level timestamps, or context biasing; Mistral's documented Voxtral Mini Transcribe 2 workflow is the more appropriate comparison in that situation. Use a speech-generation model such as Voxtral TTS when the required output is spoken audio. Use a broader multimodal or language model when the system must reason over the content, generate code, use tools, or carry on a general conversation.
Bottom line
Voxtral Mini Transcribe Realtime is a focused, low-latency transcription component rather than an all-purpose AI model. Its practical value comes from turning a continuous audio stream into text events that an application can consume immediately. At $0.006 per audio minute, it offers a clear usage-based price for this task, but teams should plan for separate components whenever the product needs speaker attribution, reasoning, tool use, or synthesized speech.

