What is Voxtral Small?
Voxtral Small is a 24-billion-parameter multimodal audio-language model provided by Mistral AI. Its main job is to understand spoken audio and respond with text or structured text. Unlike a conventional speech-to-text system that only produces a transcript, Voxtral Small can use the meaning of a recording to answer questions, summarize discussions, translate speech, and trigger functions from spoken instructions.
The model was released on July 15, 2025, and is available as an open-weight checkpoint under the Apache 2.0 license. That makes it suitable for organizations that want to investigate local deployment or build on the model directly, while Mistral also offers hosted access through its API. The hosted model identifier is voxtral-small-2507.
Within Mistral's catalog, Voxtral Small is an audio-focused model rather than a general-purpose image or video system. It combines audio input with text-language-model capabilities derived from the Mistral Small 3.1 backbone, giving it a broader role than a standalone transcription engine.
What can Voxtral Small do?
Voxtral Small accepts audio and text inputs. Audio can be supplied for transcription or interpretation, while text can provide instructions such as asking for a summary, identifying decisions in a meeting, or translating a recording. Mistral documents automatic language detection and multilingual speech understanding.
- Transcription: Convert spoken audio into text.
- Audio question answering: Ask questions about facts, topics, speakers, or events contained in a recording.
- Summarization: Condense meetings, interviews, calls, lectures, and other long-form recordings.
- Speech translation: Translate spoken content into another language through text output.
- Multilingual understanding: Process speech across supported languages with automatic language detection.
- Function calling: Convert a spoken request into a tool or function invocation, allowing an application to connect audio instructions to business actions.
- Structured responses: Return text in an application-defined structured format through the Mistral API.
Function calling is particularly useful when the model is placed between a user and an application. For example, a voice instruction such as requesting a calendar action or asking a system to retrieve a record can be interpreted as an intent and passed to an external function. The model does not independently complete an outside action; the connected application remains responsible for executing and validating the function.
Audio and context limits
Voxtral Small has a documented 32K-token context window. Context is the combined working space for the supplied instructions, transcript-related content, and generated response. A larger context can help with long recordings and detailed follow-up questions, but the practical limit also depends on how the audio is represented and on the particular API or deployment configuration.
Mistral states that Voxtral Small can handle audio of approximately 30 minutes for transcription or approximately 40 minutes for audio understanding, depending on the task and input configuration. These are provider-described task limits, not a guarantee that every recording will produce equally accurate results. Long recordings may still benefit from segmentation, preprocessing, or application-level checks.
Mistral's current model documentation does not specify a definitive knowledge-cutoff date or a maximum output-token limit for this exact model. Those values should therefore be treated as unknown rather than inferred from the context-window size. A 32K context window does not mean that the model can necessarily generate 32K output tokens.
Supported modalities and outputs
The model supports audio and text input. Its output is text, including ordinary natural-language responses, transcripts, summaries, translations, and structured text. It does not natively generate audio, images, or video.
| Capability | Voxtral Small |
|---|---|
| Audio input | Yes |
| Text input | Yes |
| Text output | Yes |
| Structured output | Supported through the Mistral API |
| Function or tool use | Yes |
| Image input or output | No |
| Video input or output | No |
| Native audio or speech output | No |
This distinction matters for voice applications. Voxtral Small can understand spoken input and produce a textual or structured response, but it is not a speech-generation or speech-to-speech model. An application that needs spoken replies would need a separate text-to-speech component.
Pricing and deployment options
Mistral's hosted API pricing is listed at $0.004 per minute of audio, plus $0.10 per million input tokens and $0.40 per million output tokens. The audio charge and token charges describe different parts of usage, so the total cost depends on recording duration, instructions, transcript-related context, and response length.
The model is also available as an open-weight checkpoint for local deployment. Local operation can provide more control over data handling and infrastructure, but Voxtral Small's 24-billion-parameter size makes it substantially more demanding to run than smaller models in the Voxtral family. The supplied research does not define a required hardware configuration, so deployment requirements should be evaluated against the specific checkpoint, quantization, serving stack, and workload.
The Apache 2.0 license is a meaningful option for teams evaluating commercial deployment, but licensing does not remove the need to assess infrastructure costs, data governance, model accuracy, and any obligations associated with the surrounding application.
Reasoning, coding, and tool use
Voxtral Small is designed primarily for audio understanding, not for advanced standalone reasoning or software development. It can follow text instructions about an audio file, compare information in a recording, organize findings, and produce structured responses. Those abilities are useful forms of task reasoning, but Mistral does not provide a definitive reasoning benchmark or a separate maximum reasoning budget for this model.
Its text capabilities come from the Mistral Small 3.1 backbone, so it can handle ordinary text instructions and code-like or structured content. However, the supplied research does not establish Voxtral Small as a specialist coding model, nor does it provide coding benchmarks. The practical reason to choose it is the combination of spoken input and text reasoning, rather than software engineering performance alone.
Tool support is more clearly defined. Voxtral Small can use function calling to represent an action requested through speech. This can support voice-driven workflows such as extracting a task from a meeting, routing a support request, searching an internal system, or passing a structured command to another service. Developers should still validate arguments and permissions outside the model, because a generated function call is an instruction for an application rather than proof that an action is safe or correct.
Main strengths and trade-offs
The model's central strength is that it brings transcription, audio interpretation, and language-model responses into one audio-to-text workflow. A separate speech-recognition step followed by a text model can work well, but it adds another component and may require managing transcript transfer, timing, errors, and integration logic. Voxtral Small is intended to handle these stages more directly.
Its multilingual audio capabilities, long-form handling, function calling, structured API responses, and open-weight availability are also useful for production systems that need more than a raw transcript. The Apache 2.0 release can be attractive to teams seeking a model they can examine or deploy outside a fully managed service.
There are important trade-offs. A 24-billion-parameter model requires more resources for local deployment than smaller alternatives. Hosted API usage is competitively priced for audio processing, but the final bill still depends on both audio minutes and token volume. The model also returns text rather than audio, so a conversational voice product needs additional speech synthesis.
For simple transcription, a specialized speech-recognition service may be easier or cheaper. For image-heavy, video, or general multimodal applications, an audio-language model is the wrong fit. For low-resource local environments, a smaller Voxtral variant or another compact speech model may be more practical, although the supplied research does not provide a direct performance comparison.
When to choose Voxtral Small
Choose Voxtral Small when audio understanding is central to the application and the system needs more than word-for-word transcription. It is a strong candidate for:
- Meeting, interview, lecture, and call summarization.
- Search and question answering over recorded audio.
- Multilingual transcription and speech translation.
- Voice-driven business workflows that require function calls.
- Applications that need structured records extracted from spoken content.
- Teams evaluating an open-weight audio model for local or controlled deployment.
It is less suitable when the main requirement is native speech generation, real-time speech-to-speech conversation, image or video understanding, or operation on very limited hardware. It may also be excessive for a basic transcription pipeline that does not need questions, summaries, translation, or actions grounded in the recording.
What to verify before production use
Before deployment, test the model with the languages, accents, recording conditions, speaker overlap, and terminology found in the target workload. Audio quality and domain vocabulary can affect transcription and downstream summaries, while a plausible answer about a recording can still contain an error.
Applications that use function calling should validate the model's selected function, arguments, permissions, and side effects. For sensitive meetings, customer calls, or regulated information, review the data-handling implications of both hosted API use and local deployment. Finally, measure end-to-end cost using actual recording lengths and response patterns rather than relying only on the per-minute audio price.
Overall, Voxtral Small is best understood as an open-weight audio-to-text reasoning model: it turns spoken content into transcripts, answers, summaries, translations, and structured actions. Its value comes from combining those tasks in one model, while its size, lack of native speech output, and unspecified maximum output limit should be part of any deployment decision.

