Codestral

Codestral Embed

by Mistral AI · Active

Codestral Embed converts source code, documentation, and software-related queries into configurable vectors for semantic search and retrieval. It supports an 8,192-token context window, dimensions up to 3,072, multiple precision formats, standard and batch API access, and pricing of $0.15 per million input tokens.

Embeddings Coding
Released on May 28, 2025, Codestral Embed is available through Mistral AI’s Embeddings API as codestral-embed-2505, with current documentation commonly using the shorter codestral-embed identifier. It accepts text such as code and developer questions, supports an 8,192-token context window, offers configurable vector dimensions up to 3,072, and provides several output precision formats for balancing retrieval quality, storage, and processing cost.
Outputs

What Codestral Embed can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

8/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Codestral
Model type Embedding
Context window 8K tokens
Release date 2025-05-28
Status Active
Knowledge cutoff notes

Mistral does not publish a model-specific knowledge cutoff for Codestral Embed. As an embedding model, its documented function is vector representation rather than conversational knowledge retrieval.

Model notes

Codestral Embed is an embedding model rather than a generative language model. The canonical dated API identifier is codestral-embed-2505, while current Mistral documentation commonly uses codestral-embed. The API has a default output dimension of 1536 and supports configurable dimensions up to 3072. Output formats include float, int8, uint8, binary, and ubinary. Mistral recommends code chunks of about 3000 characters with 1000 characters of overlap for retrieval workflows. Batch API processing is offered at a 50% discount. The editorial coding, speed, and cost scores are comparative estimates for embedding and code-retrieval workloads, not vendor benchmark scores.

Cost

Model pricing

Input $0.15 per million input tokens
Model guide

Codestral Embed: Code-Specialized Embeddings for Semantic Retrieval

Codestral Embed is Mistral AI’s embedding model for source code, documentation, and natural-language software queries. It converts these inputs into vectors for semantic code search, repository retrieval, coding-agent RAG, duplicate detection, clustering, and code analytics rather than generating text or completing code.

What is Codestral Embed?

Codestral Embed is Mistral AI’s embedding model designed specifically for software-related content. Instead of producing a written answer, it converts an input into a numerical vector, sometimes called an embedding. Texts with similar meaning are represented by vectors that are closer together, allowing an application to search for relevant material even when the wording is different.

For example, a developer might search for “where is user authentication handled?” while a repository contains a function named validate_session_token. A keyword search may not connect those phrases, but a code-specialized embedding model can help a retrieval system identify the relevant code and documentation based on meaning.

Codestral Embed can represent source code, technical documentation, issue descriptions, and natural-language queries about software. The resulting vectors can be stored in a vector database or another similarity-search system. A downstream application can then retrieve the closest matches and supply them to a separate language model, coding assistant, or analytics workflow.

Where Codestral Embed fits in Mistral AI’s catalog

Codestral Embed belongs to the Codestral family but serves a different role from a conversational or code-generation model. Its purpose is representation and retrieval: it helps software identify which code or documentation is relevant. It does not write explanations, generate source-code completions, rerank search results, moderate content, or execute tools.

The dated canonical API identifier is codestral-embed-2505. Mistral AI’s current documentation examples commonly use codestral-embed. The model is accessed through the Embeddings API rather than a chat-completions workflow, using the /v1/embeddings endpoint. Standard and batch processing are supported.

Technical specifications and vector outputs

SpecificationVerified detail
ProviderMistral AI
Release dateMay 28, 2025
Model typeCode-specialized embedding model
Context length8,192 tokens
Default vector dimension1,536
Maximum configurable dimension3,072
Output formatsfloat, int8, uint8, binary, and ubinary
Text outputNot supported; the model returns embeddings

The 8,192-token context limit determines how much input can be submitted for one embedding request. In practice, large repositories should normally be divided into smaller, meaningful pieces instead of sending entire repositories or very long files as single inputs.

The API’s default output dimension is 1,536, but applications can request a dimension up to 3,072. Mistral documents the dimensions as ordered by relevance, which means an application can retain an initial portion of a vector when it needs to reduce storage or search costs. Lower dimensions may be useful when an index contains a large number of code fragments, although the supplied research does not provide a benchmark quantifying the quality trade-off.

Codestral Embed also supports several precision formats. Float output is the conventional higher-precision choice, while int8, uint8, binary, and ubinary formats can reduce storage and processing requirements. The best choice depends on the application’s vector database, similarity-search implementation, and acceptable accuracy trade-off.

How to use it for code retrieval

A typical retrieval workflow has four stages. First, an application splits repository files, documentation, and issue records into chunks. Second, it sends those chunks to Codestral Embed and stores the returned vectors together with metadata such as file paths, programming languages, symbols, and line ranges. Third, it embeds a user’s natural-language question or code-related query. Finally, it compares the query vector with the stored vectors and returns the most similar results.

Mistral’s published guidance recommends code chunks of approximately 3,000 characters with about 1,000 characters of overlap. This is a practical starting point rather than a universal requirement. A project may need different boundaries for long functions, configuration files, generated code, documentation, or repositories with strongly interconnected modules. Preserving useful metadata and avoiding chunks that split important context can matter as much as the exact character count.

The retrieved passages can be shown directly in a repository-search interface or passed to a separate generative model for a natural-language response. Codestral Embed supplies the retrieval layer; it does not itself answer the developer’s question.

Best use cases

  • Semantic code search: Find functions, classes, configuration, or documentation by describing their purpose rather than remembering exact names.
  • Repository navigation: Help developers locate relevant implementation details across unfamiliar or large codebases.
  • Coding-agent retrieval: Supply an agent with repository context before it proposes changes, explains a bug, or prepares a patch.
  • Retrieval-augmented generation: Retrieve relevant code and documentation for a separate text-generation model to use as context.
  • Duplicate detection and similarity analysis: Identify code fragments or documents that express similar ideas despite different wording or structure.
  • Clustering and code analytics: Group related files, issues, or components for repository analysis and organization.

These uses all depend on an external indexing and similarity-search system. The model does not provide a built-in repository browser, vector database, user interface, or coding-agent runtime.

Pricing, batching, and cost trade-offs

Mistral lists standard pricing at $0.15 per million input tokens. There is no output-token price in the supplied pricing information because the model returns embeddings rather than generated text. Batch processing is available at a 50% discount, making it particularly relevant for initial repository indexing, scheduled re-indexing, or other workloads that do not require immediate results.

Codestral Embed’s cost profile is favorable for high-volume retrieval workloads compared with using a general-purpose generative model to repeatedly analyze complete files. The main operational costs are still broader than the API price: applications must store vectors, maintain indexes, manage chunking, and potentially pay for a separate model that generates the final response.

The research includes editorial scores of 8/10 for coding-related usefulness, 8/10 for speed, and 9/10 for cost. These are comparative editorial estimates for embedding and code-retrieval workloads, not scores published by Mistral AI and not standardized benchmark results.

Strengths and limitations

Codestral Embed’s clearest strength is specialization. A model intended for code and software-related language is a natural fit when the search target is a repository, technical document, issue tracker, or coding-agent context. Configurable dimensions allow teams to choose between larger vectors and more compact indexes, while quantized formats provide additional storage and processing options.

The model also fits both interactive and bulk workflows. Standard API requests can support user-facing search, while discounted batch processing can handle repository indexing. Its 8,192-token context window provides room for substantial code or documentation chunks, although sending larger chunks is not automatically better for retrieval.

Its limitations are equally important. Codestral Embed produces vectors, not text, so it cannot independently explain a search result, generate code, fix a bug, or carry out an action. It is not a reranker, meaning an application that needs a second-stage ranking system must provide one separately. It also has no documented tool use, streaming output, native image or audio input, or multimodal output in the supplied specifications.

Retrieval quality depends on choices outside the model, including chunk boundaries, metadata, similarity metrics, indexing strategy, filtering, and the quality of the downstream generation model. The supplied research does not establish a model-specific knowledge cutoff or provide benchmark results, so those details should not be assumed.

When to choose Codestral Embed

Choose Codestral Embed when the primary problem is finding or organizing software-related information by meaning. It is a strong candidate for semantic repository search, documentation retrieval, coding-agent context selection, code similarity analysis, and large-scale indexing where input cost and storage efficiency matter.

Its configurable dimensions and compact output formats are useful when an engineering team needs to manage a large vector index. Batch pricing is also relevant when indexing can happen asynchronously. The model is especially appropriate when a separate generative model already handles explanations or code changes and only a specialized retrieval component is missing.

Another option may be more appropriate when the application needs a direct conversational answer, code generation, text summarization, reranking, moderation, image or audio processing, or tool execution. In those cases, Codestral Embed can still be one component of a larger pipeline, but it should not be treated as a complete coding assistant or general-purpose language model.

Bottom line

Codestral Embed is a focused infrastructure model rather than an end-user chatbot. It turns code and natural-language software queries into searchable vectors, with an 8,192-token context window, dimensions up to 3,072, multiple precision formats, and discounted batch processing. Its value comes from improving retrieval and organization around code; applications must supply the index, search logic, interface, and any model responsible for the final written response.


Answers to Frequently Asked Questions

How much does Codestral Embed cost?
Mistral lists standard pricing at $0.15 per million input tokens. Batch processing is available at a 50% discount, making it useful for bulk repository indexing and scheduled re-indexing. Applications must also account for vector storage, index maintenance, and any separate model used to generate final responses.
What are the main technical specifications of Codestral Embed?
Codestral Embed has an 8,192-token context length, a default vector dimension of 1,536, and a configurable maximum dimension of 3,072. It supports float, int8, uint8, binary, and ubinary output formats, and is accessed through the Embeddings API.
Does Codestral Embed generate code or answer developer questions?
No. Codestral Embed is an embedding model that returns vectors rather than text. It does not generate code, explain search results, fix bugs, rerank results, or execute tools. A separate search system and, if needed, a generative language model must handle those tasks.
What is Codestral Embed used for?
Codestral Embed converts source code, technical documentation, issue descriptions, and software-related queries into numerical vectors for semantic search. It can support repository navigation, coding-agent retrieval, retrieval-augmented generation, duplicate detection, clustering, and code similarity analysis.


Sources 6
Provider

About Mistral AI