What is the Mistral API and when should you use it?
The Mistral API is a hosted developer interface for sending requests to Mistral models and receiving generated results. Mistral Studio provides the surrounding workspace for API keys, model experimentation, the Playground, usage monitoring, organization controls, and related developer services. The main global API endpoint is https://api.mistral.ai/v1.
The simplest starting point is the Chat Completions API. It accepts a model name and a sequence of messages, then returns a model response. This approach is suitable for conversational applications, text generation, multimodal requests on supported models, streaming, structured output, and conventional application-managed tool calling.
Use the Agents and Conversations APIs when you need reusable instructions, persistent developer-managed state, stored interaction history, built-in tools, or multi-step workflows. Agents define reusable configurations containing a model, instructions, tools, completion parameters, and optional guardrails. Conversations can use an agent or a base model and can be configured with store=false when automatic storage of new conversation history is not wanted.
Who is the API for?
The API is intended for developers building chat applications, retrieval-augmented generation systems, document-processing workflows, coding tools, internal assistants, autonomous or semi-autonomous agents, and other software that needs model inference. It is also relevant to organizations that need model access, regional processing options, usage controls, or private and enterprise deployment choices.
It is less suitable if your project requires unrestricted high-volume usage without quotas, consistently deterministic answers, native consumer video generation, or a mature plugin marketplace. Model responses remain probabilistic, so important outputs and tool arguments require validation.
Getting access and creating an API key
Create or access a Mistral Studio account, then create an API key in Studio or the administrative API-key area. Applications normally send the key as a Bearer token in the HTTP Authorization header. Store it in an environment variable or a managed secret store rather than committing it to source code.
The examples below use the environment variable MISTRAL_API_KEY. The key is required for API requests, and the API uses usage-based billing rather than a single universal subscription price.
Choosing an API and model
For a first integration, choose a currently available chat model and use Chat Completions. Model identifiers, availability, pricing, and lifecycle status can change, so check the live Mistral documentation and pricing pages before deploying. A model alias is convenient when automatic model updates are acceptable. Use a dated model identifier when reproducibility is more important than automatic upgrades.
| Requirement | Recommended starting point |
|---|---|
| One request and one response | Chat Completions |
| Visible output as it is generated | Chat Completions with streaming |
| Reliable machine-readable fields | Structured outputs or JSON Schema on a compatible model and endpoint |
| Your application must execute an external operation | User-defined function calling |
| Reusable instructions and stored workflow state | Agents and Conversations |
| Document extraction or question answering | Document AI and OCR |
| Semantic search or retrieval | Embeddings, Files, and Libraries |
Vision-capable models can accept image URLs or base64-encoded images through Chat Completions. Check model compatibility before relying on vision or other specialized features.
Making your first API request
The standard request is a POST to /v1/chat/completions. It includes a model and an array of messages. The response normally places generated text at choices[0].message.content.
First request with cURL
curl --fail-with-body --silent --show-error https://api.mistral.ai/v1/chat/completions
-H "Authorization: Bearer ${MISTRAL_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "mistral-large-latest",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain API rate limits in one paragraph."}
],
"temperature": 0.2,
"max_tokens": 200
}' | jq -r '.choices[0].message.content'The system message sets high-level behavior, while the user message contains the task. Parameters such as temperature and max_tokens affect generation behavior and output length. Use only parameters supported by the selected model and current API documentation.
First request with the Python SDK
# Install: pip install -U mistralai
import os
from mistralai.client import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
model="mistral-large-latest",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Return two practical uses of an API gateway."},
],
temperature=0.2,
)
print(response.choices[0].message.content)Understanding the response
A normal chat response contains a list of choices. For a basic request, the first choice's message content is the generated answer. Applications using tools or structured output should inspect the corresponding response fields rather than assuming that every response is plain text.
When a model returns a function call, the application must read the function name and arguments, validate them, execute the approved local or external operation, and send the result back in a follow-up request. The model does not automatically grant permission to perform the operation. Treat generated arguments as untrusted input.
How Mistral API pricing works
Mistral API pricing is primarily usage-based. Text models generally charge separately for input tokens and output tokens, with rates stated per million tokens. The exact amount depends on the selected model and can change independently of the API surface.
Other services use different billing units. OCR is priced per thousand pages, speech services can be priced per audio minute, and some built-in tools use per-call or per-thousand-call pricing. Batch processing is listed at approximately half of standard pricing for eligible asynchronous workloads. Cached input can receive substantial discounts, and the research identifies discounts of up to 90% for eligible repeated prompts. Priority inference is intended for workloads that need more predictable or faster capacity.
Regional inference for supported workloads uses EU or US endpoints and adds a 10% list-price surcharge. Before estimating production costs, check the live Mistral API pricing page for the exact model, service, region, and billing unit you plan to use.
What can developers build with the API?
Streaming responses
Chat Completions supports server-sent event streaming with stream=true. Instead of waiting for the complete answer, your application receives incremental events and can display text as it arrives. Streaming usually improves time to first visible output, but it does not necessarily reduce total generation time.
curl --fail-with-body --silent --show-error https://api.mistral.ai/v1/chat/completions
-H "Authorization: Bearer ${MISTRAL_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "mistral-large-latest",
"messages": [{"role": "user", "content": "Give three uses for structured outputs."}],
"stream": true
}'Function calling and built-in tools
Function calling lets a model request an operation described by your application. For example, you can provide a get_weather function with a city argument. Your application remains responsible for checking the arguments, calling the weather service, handling errors, and returning the result to the model.
Agents also support built-in tools, including web search, premium web search, code execution, image generation, and document-library search. Web search is priced separately from model-token usage. Regional inference currently has important restrictions: stateful Agents, Batch, and Files API features are not available on regional endpoints, and function calling is the only regional tool identified in the supplied documentation.
Structured outputs
Compatible requests can use JSON mode or JSON Schema mode through the response-format configuration. Structured output is useful when a downstream program needs predictable fields, such as a classification result, an extracted invoice object, or a list of actions. JSON Schema is preferable when the application needs a defined structure, but the selected model and endpoint must support the requested mode.
Vision, documents, files, and libraries
Selected vision-capable models accept images through URLs or base64-encoded data. Document AI and OCR support extraction, annotations, structured output, and document question answering. Files can be uploaded and indexed into persistent Libraries, which agents can search for retrieval-augmented workflows.
Files and libraries are stateful services: they may require storage to provide their functionality. Do not assume that a zero-data-retention setting automatically covers uploaded documents, libraries, agents, or every other API feature.
Embeddings and related services
Embedding models convert text into numerical representations used for semantic search and retrieval. The broader platform also includes moderation, audio services, image generation, batch processing, regional inference, and administrative usage APIs. Availability and pricing are service- and model-dependent.
Using Agents and Conversations for persistent workflows
Chat Completions is generally the clearest choice when your application owns the conversation history and tool orchestration. Agents and Conversations are more appropriate when you want Mistral's platform to manage reusable agent configuration, persistent interaction state, built-in tools, or handoffs between agents.
A conversation can be created with an agent or a base model. If you do not want automatic cloud storage of new conversation history, the API supports store=false for conversation creation. This setting does not make every related service stateless and should be evaluated separately from organization-level retention controls.
SDKs, Playground, and developer tools
Mistral officially supports Python and TypeScript SDKs. The Python package is mistralai; the TypeScript or JavaScript package is @mistralai/mistralai. Other languages can call the REST API directly, and third-party libraries may be available, but they are not equivalent to first-party SDK support.
Mistral Studio includes a Playground for experimenting with models and prompts before writing application code. It also provides API-key management, organization controls, usage monitoring, limits information, and related administrative tools. For coding workflows, the wider Mistral ecosystem includes CLI, VS Code, and remote web-session options through Vibe Code, but those consumer and coding products are separate from the core API integration described here.
Limits, privacy, and production considerations
Usage limits and latency
Limits vary by organization, plan, model, API, and region. They can include requests per second, tokens per minute, audio seconds per minute, OCR pages per minute, document-upload limits, and monthly usage. Check the limits area in Studio for the limits that apply to your account. Higher limits can be requested by describing expected request rate, token throughput, monthly volume, and use case.
Latency depends on model size, prompt and output length, queue conditions, region, and service tier. Streaming improves the time before the first visible output. Priority inference is intended for more predictable capacity and faster service, while regional inference can address geographic processing requirements when the selected model and feature are supported.
Privacy, training, and retention
Data handling depends on the applicable plan, contract, product, API feature, and organization settings. Eligible paid organizations can request or enable zero-data-retention controls for supported stateless API calls. When enabled, covered inputs and outputs are not stored or logged longer than required to generate the response.
Zero-data-retention does not automatically cover stateful Agents, Conversations, Libraries, uploaded documents, operational metadata, account settings, billing records, or usage analytics. Review Mistral's current commercial terms, privacy documentation, and organization configuration before sending sensitive data.
Reliability and safety practices
- Keep API keys out of source control and use a managed secret store in production.
- Implement timeouts, retries, and exponential backoff for transient failures and rate-limit responses.
- Set explicit output limits and monitor token usage.
- Validate every model-generated tool argument before executing it.
- Use structured outputs when downstream systems require predictable fields, then validate the returned data anyway.
- Pin a dated model identifier when reproducibility matters, and monitor model lifecycle notices.
- Confirm regional endpoint support before depending on Agents, Batch, Files, or other stateful features.
- Review retention and training settings for each service rather than applying one policy assumption to the whole platform.
When is the Mistral API a good or poor choice?
The Mistral API is a good choice when you want a familiar chat-completions interface, official Python and TypeScript SDKs, fast access to Mistral models, multilingual capabilities, multimodal input, document processing, structured output, function calling, embeddings, or a path from simple requests to persistent agents and built-in tools. Regional inference and enterprise-oriented controls can also matter when processing location and governance are requirements.
It may be a poor fit when your application depends on fixed pricing independent of usage, unrestricted throughput, guaranteed deterministic responses, universal regional support for every feature, or a mature consumer plugin marketplace. Fine-tuning also requires particular care: legacy fine-tuning documentation is marked deprecated, while current pricing lists selected classifier fine-tuning products. Verify that the intended model and workflow are currently supported before designing around fine-tuning.
