Developer platform

Mistral API

Developer overview for Mistral AI, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Chat Completions API
SDK support Official Python SDK package mistralai and official TypeScript/JavaScript SDK package @mistralai/mistralai. Third-party SDKs are available for other languages.
Rate limits Limits vary by organization, plan, model, API, and region. Documented controls can include requests per second, tokens per minute, audio seconds per minute, OCR pages per minute, document upload limits, and monthly usage. Higher limits can be requested fr
Platform

API overview

Endpoints

API access

Base URL https://api.mistral.ai/v1
Primary API Chat Completions API
Pricing

API pricing

Pricing model Usage-based pricing by model and service; token-based for most text models, with separate pricing units for OCR, audio, tools, batch, priority, cached input, and fine-tuning products.

Text models are generally priced per million input and output tokens. Batch processing is listed at 50% of standard pricing, cached input can receive discounts of up to 90%, and regional inference adds 10%. Current prices vary by model and service.

Developer experience

SDKs & usability

SDKs Official Python SDK package mistralai and official TypeScript/JavaScript SDK package @mistralai/mistralai. Third-party SDKs are available for other languages.
Ease of use High for standard chat requests because the REST API follows a familiar chat-completions structure and official Python and TypeScript SDKs are available. Advanced agents, libraries, regional routing, and tool workflows require additional concepts.
Documentation Good and broad, with current quickstarts, API reference pages, SDK documentation, cookbooks, model lifecycle information, regional-inference guidance, and explicit deprecation notices. Some legacy fine-tuning material remains marked deprecated.
Latency Variable according to model, prompt length, output length, queue conditions, region, and service tier. Streaming improves time to first visible token. Priority inference is intended for more predictable capacity and faster service. Regional inference can
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Fine-tuning
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

Mistral Studio is the current developer platform. Chat Completions is the recommended general-purpose starting point for stateless text and multimodal requests. The platform also provides OCR and Document AI, embeddings, Files and Libraries, Agents and Conversations, built-in web search, code execution, image generation, moderation, batch processing, regional inference, and administrative usage APIs. Vision is available on selected models through Chat Completions using image URLs or base64 data. Structured JSON and JSON Schema outputs are supported on compatible models and endpoints. Tool calling can be user-defined function calling or built-in agent tools. Agents and Conversations provide persistent developer-managed state, but regional inference endpoints do not support stateful Agents, Batch, or Files API features. Legacy fine-tuning documentation is explicitly deprecated; current pricing lists selected classifier fine-tuning products, so fine-tuning must be verified for the intended model and workflow. Mistral documents standard, batch, priority, cached-input, and regional inference options. The global endpoint is not region-specific, while EU and US endpoints provide regional inference for supported models. Use retries with exponential backoff for transient failures and rate limits, and validate all model-generated tool arguments before execution.

Examples

API examples

curl --fail-with-body --silent --show-error https://api.mistral.ai/v1/chat/completions \
  -H "Authorization: Bearer ${MISTRAL_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-large-latest",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain API rate limits in one paragraph."}
    ],
    "temperature": 0.2,
    "max_tokens": 200
  }' | jq -r '.choices[0].message.content'

# Streaming example
curl --fail-with-body --silent --show-error https://api.mistral.ai/v1/chat/completions \
  -H "Authorization: Bearer ${MISTRAL_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-large-latest",
    "messages": [{"role": "user", "content": "Give three uses for structured outputs."}],
    "stream": true
  }'
# Install: pip install -U mistralai
import json
import os
from mistralai.client import Mistral

api_key = os.environ["MISTRAL_API_KEY"]
client = Mistral(api_key=api_key)

response = client.chat.complete(
    model="mistral-large-latest",
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "Return two practical uses of an API gateway."},
    ],
    temperature=0.2,
)
print(response.choices[0].message.content)

# Streaming
for event in client.chat.stream(
    model="mistral-large-latest",
    messages=[{"role": "user", "content": "Explain retries with exponential backoff."}],
):
    if event.data.choices and event.data.choices[0].delta.content:
        print(event.data.choices[0].delta.content, end="", flush=True)
print()

# Function calling
weather_tool = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"]
        }
    }
}]

tool_response = client.chat.complete(
    model="mistral-medium-latest",
    messages=[{"role": "user", "content": "What is the weather in Paris?"}],
    tools=weather_tool,
)
message = tool_response.choices[0].message
if message.tool_calls:
    call = message.tool_calls[0]
    arguments = json.loads(call.function.arguments)
    print({"tool": call.function.name, "arguments": arguments})
// Install: npm install @mistralai/mistralai
import { Mistral } from "@mistralai/mistralai";

const apiKey = process.env.MISTRAL_API_KEY;
if (!apiKey) throw new Error("MISTRAL_API_KEY is not set");

const client = new Mistral({ apiKey });

try {
  const response = await client.chat.complete({
    model: "mistral-large-latest",
    messages: [
      { role: "system", content: "You are a concise technical assistant." },
      { role: "user", content: "Explain why API clients should use timeouts." }
    ],
    temperature: 0.2,
    maxTokens: 200
  });
  console.log(response.choices[0].message.content);

  const toolResponse = await client.chat.complete({
    model: "mistral-medium-latest",
    messages: [{ role: "user", content: "What is the weather in Paris?" }],
    tools: [{
      type: "function",
      function: {
        name: "get_weather",
        description: "Get current weather for a city.",
        parameters: {
          type: "object",
          properties: { city: { type: "string" } },
          required: ["city"]
        }
      }
    }]
  });

  const toolCalls = toolResponse.choices[0].message.toolCalls ?? [];
  for (const call of toolCalls) {
    console.log(call.function.name, JSON.parse(call.function.arguments));
  }
} catch (error) {
  console.error(error);
  process.exitCode = 1;
}
<?php
$apiKey = getenv('MISTRAL_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('MISTRAL_API_KEY is not set');
}

$url = 'https://api.mistral.ai/v1/chat/completions';
$payload = [
    'model' => 'mistral-large-latest',
    'messages' => [
        ['role' => 'system', 'content' => 'You are a concise technical assistant.'],
        ['role' => 'user', 'content' => 'Explain the purpose of structured JSON output.']
    ],
    'temperature' => 0.2,
    'max_tokens' => 200
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60
]);

$body = curl_exec($ch);
if ($body === false) {
    throw new RuntimeException('cURL error: ' . curl_error($ch));
}
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true);
if ($status < 200 || $status >= 300) {
    $message = $data['message'] ?? $body;
    throw new RuntimeException('Mistral API HTTP ' . $status . ': ' . $message);
}

$content = $data['choices'][0]['message']['content'] ?? null;
if ($content === null) {
    throw new RuntimeException('The response did not contain message content');
}
echo $content . PHP_EOL;
?>
Policies

Data & usage

Data training

Mistral's exact data-use terms depend on the applicable plan, contract, product, and API feature. API customers should review the current commercial terms and privacy documentation rather than assume that all Studio features have identical data handling. Paid organizations can request or enable zero-data-retention controls for supported stateless API calls. Stateful Agents, Conversations, Libraries, and uploaded documents may require storage to provide their functionality.

Data retention

Zero data retention is available for eligible paid organizations and supported stateless API endpoints. When enabled, covered inputs and outputs are not stored or logged longer than required to generate the response. ZDR does not automatically cover all stateful APIs, files, libraries, agents, operational metadata, account settings, billing records, or usage analytics. Conversation creation can use store=false to avoid automatic storage of new conversation history.

Rate limits

Limits vary by organization, plan, model, API, and region. Documented controls can include requests per second, tokens per minute, audio seconds per minute, OCR pages per minute, document upload limits, and monthly usage. Higher limits can be requested fr

Developer guide

Mistral AI API Guide: Studio, Models, Agents, Pricing, and SDKs

Mistral Studio and the Mistral API provide hosted access to Mistral models through REST endpoints and official Python and TypeScript SDKs. Developers can build chat and multimodal applications, stream responses, request structured JSON, call external functions, process documents, generate embeddings, create persistent agents and conversations, search libraries, and use built-in tools such as web search and code execution.
Mistral AI's developer platform is centered on Mistral Studio and the Mistral API. A beginner can start with a single authenticated Chat Completions request, while larger applications can use agents, conversations, document processing, libraries, regional inference, and administrative usage controls. This guide explains how access works, how to make a first request, what the platform supports, how pricing is organized, and which limitations matter in production.
Mistral Studio provides REST APIs and official Python and TypeScript SDKs for chat, multimodal inference, streaming, structured outputs, function calling, agents, conversations, OCR, embeddings, document libraries, and web search. Pricing is usage-based, while limits, regional support, retention, and model availability vary by service.

What is the Mistral API and when should you use it?

The Mistral API is a hosted developer interface for sending requests to Mistral models and receiving generated results. Mistral Studio provides the surrounding workspace for API keys, model experimentation, the Playground, usage monitoring, organization controls, and related developer services. The main global API endpoint is https://api.mistral.ai/v1.

The simplest starting point is the Chat Completions API. It accepts a model name and a sequence of messages, then returns a model response. This approach is suitable for conversational applications, text generation, multimodal requests on supported models, streaming, structured output, and conventional application-managed tool calling.

Use the Agents and Conversations APIs when you need reusable instructions, persistent developer-managed state, stored interaction history, built-in tools, or multi-step workflows. Agents define reusable configurations containing a model, instructions, tools, completion parameters, and optional guardrails. Conversations can use an agent or a base model and can be configured with store=false when automatic storage of new conversation history is not wanted.

Who is the API for?

The API is intended for developers building chat applications, retrieval-augmented generation systems, document-processing workflows, coding tools, internal assistants, autonomous or semi-autonomous agents, and other software that needs model inference. It is also relevant to organizations that need model access, regional processing options, usage controls, or private and enterprise deployment choices.

It is less suitable if your project requires unrestricted high-volume usage without quotas, consistently deterministic answers, native consumer video generation, or a mature plugin marketplace. Model responses remain probabilistic, so important outputs and tool arguments require validation.

Getting access and creating an API key

Create or access a Mistral Studio account, then create an API key in Studio or the administrative API-key area. Applications normally send the key as a Bearer token in the HTTP Authorization header. Store it in an environment variable or a managed secret store rather than committing it to source code.

The examples below use the environment variable MISTRAL_API_KEY. The key is required for API requests, and the API uses usage-based billing rather than a single universal subscription price.

Choosing an API and model

For a first integration, choose a currently available chat model and use Chat Completions. Model identifiers, availability, pricing, and lifecycle status can change, so check the live Mistral documentation and pricing pages before deploying. A model alias is convenient when automatic model updates are acceptable. Use a dated model identifier when reproducibility is more important than automatic upgrades.

RequirementRecommended starting point
One request and one responseChat Completions
Visible output as it is generatedChat Completions with streaming
Reliable machine-readable fieldsStructured outputs or JSON Schema on a compatible model and endpoint
Your application must execute an external operationUser-defined function calling
Reusable instructions and stored workflow stateAgents and Conversations
Document extraction or question answeringDocument AI and OCR
Semantic search or retrievalEmbeddings, Files, and Libraries

Vision-capable models can accept image URLs or base64-encoded images through Chat Completions. Check model compatibility before relying on vision or other specialized features.

Making your first API request

The standard request is a POST to /v1/chat/completions. It includes a model and an array of messages. The response normally places generated text at choices[0].message.content.

First request with cURL

curl --fail-with-body --silent --show-error https://api.mistral.ai/v1/chat/completions 
  -H "Authorization: Bearer ${MISTRAL_API_KEY}" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "mistral-large-latest",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain API rate limits in one paragraph."}
    ],
    "temperature": 0.2,
    "max_tokens": 200
  }' | jq -r '.choices[0].message.content'

The system message sets high-level behavior, while the user message contains the task. Parameters such as temperature and max_tokens affect generation behavior and output length. Use only parameters supported by the selected model and current API documentation.

First request with the Python SDK

# Install: pip install -U mistralai
import os
from mistralai.client import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

response = client.chat.complete(
    model="mistral-large-latest",
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "Return two practical uses of an API gateway."},
    ],
    temperature=0.2,
)

print(response.choices[0].message.content)

Understanding the response

A normal chat response contains a list of choices. For a basic request, the first choice's message content is the generated answer. Applications using tools or structured output should inspect the corresponding response fields rather than assuming that every response is plain text.

When a model returns a function call, the application must read the function name and arguments, validate them, execute the approved local or external operation, and send the result back in a follow-up request. The model does not automatically grant permission to perform the operation. Treat generated arguments as untrusted input.

How Mistral API pricing works

Mistral API pricing is primarily usage-based. Text models generally charge separately for input tokens and output tokens, with rates stated per million tokens. The exact amount depends on the selected model and can change independently of the API surface.

Other services use different billing units. OCR is priced per thousand pages, speech services can be priced per audio minute, and some built-in tools use per-call or per-thousand-call pricing. Batch processing is listed at approximately half of standard pricing for eligible asynchronous workloads. Cached input can receive substantial discounts, and the research identifies discounts of up to 90% for eligible repeated prompts. Priority inference is intended for workloads that need more predictable or faster capacity.

Regional inference for supported workloads uses EU or US endpoints and adds a 10% list-price surcharge. Before estimating production costs, check the live Mistral API pricing page for the exact model, service, region, and billing unit you plan to use.

What can developers build with the API?

Streaming responses

Chat Completions supports server-sent event streaming with stream=true. Instead of waiting for the complete answer, your application receives incremental events and can display text as it arrives. Streaming usually improves time to first visible output, but it does not necessarily reduce total generation time.

curl --fail-with-body --silent --show-error https://api.mistral.ai/v1/chat/completions 
  -H "Authorization: Bearer ${MISTRAL_API_KEY}" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "mistral-large-latest",
    "messages": [{"role": "user", "content": "Give three uses for structured outputs."}],
    "stream": true
  }'

Function calling and built-in tools

Function calling lets a model request an operation described by your application. For example, you can provide a get_weather function with a city argument. Your application remains responsible for checking the arguments, calling the weather service, handling errors, and returning the result to the model.

Agents also support built-in tools, including web search, premium web search, code execution, image generation, and document-library search. Web search is priced separately from model-token usage. Regional inference currently has important restrictions: stateful Agents, Batch, and Files API features are not available on regional endpoints, and function calling is the only regional tool identified in the supplied documentation.

Structured outputs

Compatible requests can use JSON mode or JSON Schema mode through the response-format configuration. Structured output is useful when a downstream program needs predictable fields, such as a classification result, an extracted invoice object, or a list of actions. JSON Schema is preferable when the application needs a defined structure, but the selected model and endpoint must support the requested mode.

Vision, documents, files, and libraries

Selected vision-capable models accept images through URLs or base64-encoded data. Document AI and OCR support extraction, annotations, structured output, and document question answering. Files can be uploaded and indexed into persistent Libraries, which agents can search for retrieval-augmented workflows.

Files and libraries are stateful services: they may require storage to provide their functionality. Do not assume that a zero-data-retention setting automatically covers uploaded documents, libraries, agents, or every other API feature.

Embeddings and related services

Embedding models convert text into numerical representations used for semantic search and retrieval. The broader platform also includes moderation, audio services, image generation, batch processing, regional inference, and administrative usage APIs. Availability and pricing are service- and model-dependent.

Using Agents and Conversations for persistent workflows

Chat Completions is generally the clearest choice when your application owns the conversation history and tool orchestration. Agents and Conversations are more appropriate when you want Mistral's platform to manage reusable agent configuration, persistent interaction state, built-in tools, or handoffs between agents.

A conversation can be created with an agent or a base model. If you do not want automatic cloud storage of new conversation history, the API supports store=false for conversation creation. This setting does not make every related service stateless and should be evaluated separately from organization-level retention controls.

SDKs, Playground, and developer tools

Mistral officially supports Python and TypeScript SDKs. The Python package is mistralai; the TypeScript or JavaScript package is @mistralai/mistralai. Other languages can call the REST API directly, and third-party libraries may be available, but they are not equivalent to first-party SDK support.

Mistral Studio includes a Playground for experimenting with models and prompts before writing application code. It also provides API-key management, organization controls, usage monitoring, limits information, and related administrative tools. For coding workflows, the wider Mistral ecosystem includes CLI, VS Code, and remote web-session options through Vibe Code, but those consumer and coding products are separate from the core API integration described here.

Limits, privacy, and production considerations

Usage limits and latency

Limits vary by organization, plan, model, API, and region. They can include requests per second, tokens per minute, audio seconds per minute, OCR pages per minute, document-upload limits, and monthly usage. Check the limits area in Studio for the limits that apply to your account. Higher limits can be requested by describing expected request rate, token throughput, monthly volume, and use case.

Latency depends on model size, prompt and output length, queue conditions, region, and service tier. Streaming improves the time before the first visible output. Priority inference is intended for more predictable capacity and faster service, while regional inference can address geographic processing requirements when the selected model and feature are supported.

Privacy, training, and retention

Data handling depends on the applicable plan, contract, product, API feature, and organization settings. Eligible paid organizations can request or enable zero-data-retention controls for supported stateless API calls. When enabled, covered inputs and outputs are not stored or logged longer than required to generate the response.

Zero-data-retention does not automatically cover stateful Agents, Conversations, Libraries, uploaded documents, operational metadata, account settings, billing records, or usage analytics. Review Mistral's current commercial terms, privacy documentation, and organization configuration before sending sensitive data.

Reliability and safety practices

  1. Keep API keys out of source control and use a managed secret store in production.
  2. Implement timeouts, retries, and exponential backoff for transient failures and rate-limit responses.
  3. Set explicit output limits and monitor token usage.
  4. Validate every model-generated tool argument before executing it.
  5. Use structured outputs when downstream systems require predictable fields, then validate the returned data anyway.
  6. Pin a dated model identifier when reproducibility matters, and monitor model lifecycle notices.
  7. Confirm regional endpoint support before depending on Agents, Batch, Files, or other stateful features.
  8. Review retention and training settings for each service rather than applying one policy assumption to the whole platform.

When is the Mistral API a good or poor choice?

The Mistral API is a good choice when you want a familiar chat-completions interface, official Python and TypeScript SDKs, fast access to Mistral models, multilingual capabilities, multimodal input, document processing, structured output, function calling, embeddings, or a path from simple requests to persistent agents and built-in tools. Regional inference and enterprise-oriented controls can also matter when processing location and governance are requirements.

It may be a poor fit when your application depends on fixed pricing independent of usage, unrestricted throughput, guaranteed deterministic responses, universal regional support for every feature, or a mature consumer plugin marketplace. Fine-tuning also requires particular care: legacy fine-tuning documentation is marked deprecated, while current pricing lists selected classifier fine-tuning products. Verify that the intended model and workflow are currently supported before designing around fine-tuning.

Sources 20