Embed v3.0

embed-multilingual-v3.0

by Cohere · Active

Cohere embed-multilingual-v3.0 is a 1,024-dimensional embedding model for multilingual semantic search, cross-lingual retrieval, classification, clustering, and image-text similarity. It supports more than 100 languages, accepts text and images, and has a 512-token input limit. The model is designed for retrieval pipelines rather than chat or generation, with pricing requiring confirmation for current accounts and deployments.

Embeddings Reasoning Coding
Cohere embed-multilingual-v3.0 is a specialized embedding model for applications that need to compare the meaning of content across languages. It accepts text and image inputs, produces 1,024-dimensional vectors, and supports more than 100 languages. Its 512-token input limit and lack of generative output make it a focused choice for retrieval, indexing, classification, clustering, and similarity workflows rather than chat or content generation.
Outputs

What embed-multilingual-v3.0 can produce

Embeddings
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Embed v3.0
Model type Other
Context window 512 tokens
Release date 2023-11-02
Status Active
Knowledge cutoff notes

Cohere does not publish a direct knowledge-cutoff date for this embedding model. Embedding models are used for vector representation rather than open-ended factual generation.

Model notes

Canonical API model ID is embed-multilingual-v3.0. Cohere documents support for more than 100 languages, fixed 1,024-dimensional embeddings, and a 512-token maximum input length. Embed v3 models gained image embedding support on October 22, 2024; image requests accept PNG, JPEG, WebP, or GIF data URLs, allow one image per request, and impose a 5 MB image limit. Text requests can use search_document, search_query, classification, or clustering input types. Embed Jobs provides asynchronous batch processing. The public current Cohere pricing page does not display an exact self-serve price for this legacy model; $0.10 per 1M input tokens is reported by third-party model-pricing catalogs and should be confirmed against the applicable Cohere account or deployment. Cohere's newer embed-v4.0 is a successor for many workloads but is not the same model.

Cost

Model pricing

Input $0.10 per 1M input tokens
Model guide

Cohere Embed Multilingual v3.0 for Cross-Lingual Search and Image Retrieval

Cohere embed-multilingual-v3.0 is a multilingual embedding model that converts text and images into fixed 1,024-dimensional vectors for semantic search, cross-lingual retrieval, classification, clustering, and image-text similarity across more than 100 languages.

What is Cohere embed-multilingual-v3.0?

Cohere embed-multilingual-v3.0 is an embedding model from Cohere. Instead of writing an answer in natural language, it converts content into a numerical representation called an embedding. Texts with similar meanings should produce vectors that are close together when measured with a suitable similarity method.

This makes the model useful for systems that need to find, group, classify, or compare content. For example, a search application can embed a user's query in one language and compare it with documents written in another. A support platform can also use embeddings to classify incoming messages, while a recommendation system can compare descriptions, queries, or images.

The model is part of Cohere's Embed family and is specifically oriented toward multilingual workloads. It is not a general-purpose language model and does not generate prose, hold conversations, execute code, or call tools.

Key specifications at a glance

SpecificationDetails
ProviderCohere
Canonical model IDembed-multilingual-v3.0
Model typeMultilingual embedding model
Release dateNovember 2, 2023
Supported languagesMore than 100
Embedding size1,024 dimensions
Maximum input length512 tokens
Text inputSupported
Image inputSupported for Embed v3 models
Text, image, audio, and video outputNot supported; the output is an embedding vector
Primary accessCohere Embed and Embed Jobs endpoints, with selected cloud deployments

The dimensions are fixed: every successful embedding has 1,024 numerical components. The model documentation lists cosine similarity, dot product similarity, and Euclidean distance as supported ways to compare vectors.

How multilingual search works with the model

In a conventional keyword search system, a query and a document often need to share the same words. Embedding search instead compares representations of meaning. With a multilingual embedding model, a query in one language can be compared with documents in another language when both are represented in the same vector space.

A typical retrieval workflow has two stages. First, an organization embeds its documents and stores the resulting vectors in a vector database or another search index. When a user submits a query, the application embeds the query and retrieves documents whose vectors are most similar. A separate generative model may then use those retrieved documents to produce an answer, but embed-multilingual-v3.0 itself only performs the representation and matching step.

The model can therefore support multilingual knowledge bases, cross-language help centers, product catalogs, internal search, and retrieval-augmented generation pipelines. The research supports more than 100 languages, including English, Spanish, French, Chinese, Arabic, Japanese, Korean, and Hindi.

Supported inputs and task-specific modes

For text, Cohere provides task-specific input types that tell the embedding endpoint how the vector will be used. The documented modes include search_document, search_query, classification, and clustering.

  • search_document: Use when embedding documents, articles, product records, or other content that will be searched.
  • search_query: Use for a user's search request or another query intended to retrieve documents.
  • classification: Use when vectors will support categorization or labeling.
  • clustering: Use when grouping similar items without necessarily having predefined labels.

These modes are more specific than sending every piece of text through one undifferentiated embedding setting. For search systems, using the document and query modes for their respective data types is the relevant configuration to evaluate.

Embed v3 models also gained image embedding support in an October 2024 update. Image requests use the image input type and accept PNG, JPEG, WebP, or GIF images supplied as base64 data URLs. The documented restriction is one image per request, with a maximum image size of 5 MB. This enables image-to-text and text-to-image retrieval when the application embeds both modalities and compares their vectors.

Best use cases

Multilingual semantic search: Search across documents written in multiple languages without requiring every document to be translated first.

Cross-lingual retrieval: Return relevant Spanish, Arabic, Japanese, or other language content for a query written in English, subject to the quality of the language and domain coverage.

RAG document indexing: Create the retrieval layer for a retrieval-augmented generation system. The model can index source material and match user questions to relevant passages before a separate generation model formulates a response.

Classification: Represent support tickets, reviews, documents, or messages for downstream classification workflows.

Clustering: Group multilingual content by semantic similarity, such as organizing customer feedback or discovering themes in a large document collection.

Similarity and recommendation systems: Compare descriptions, records, or other content to identify related items.

Image-text retrieval: Use the image capability to connect visual assets with text queries or compare images with one another in a shared retrieval workflow.

Important limitations and trade-offs

The main limitation is the 512-token maximum input length. Long documents must be split into smaller passages before embedding. That chunking step can affect retrieval quality: very small chunks may lose context, while very large chunks may exceed the model's limit or contain too many unrelated ideas. The supplied research does not specify a single best chunking strategy, so applications should test chunk size and overlap on their own data.

The output is a vector, not an answer. The model cannot summarize a retrieved passage, explain a result, write code, or conduct a conversation. It also has no documented tool or function-calling capability, no streaming output, and no reasoning mode in the sense used for generative reasoning models.

Image support is useful but constrained. Each image request is limited to one image, images must be supplied in an accepted format as base64 data URLs, and each image may be no larger than 5 MB. The model should not be treated as an image generator or an image-understanding chat assistant.

Its fixed 1,024-dimensional output may be a practical advantage for predictable index sizing, but it does not offer the configurable output dimensions documented for Cohere's newer embed-v4.0. Embed v4.0 also provides a larger context window and mixed text-image inputs, so it may be more suitable for new workloads that need those capabilities. That comparison does not make v3.0 obsolete for every use case: v3.0 remains a distinct model with multilingual support and an independently usable API.

Speed, cost, and pricing

Embedding models are generally used in high-volume indexing and retrieval pipelines, where cost and throughput can matter more than generative response quality. The supplied editorial evaluation rates this model highly for speed and cost efficiency relative to the alternatives considered, but those are assessment scores rather than Cohere-published benchmarks and should not be treated as official performance guarantees.

A third-party model-pricing catalog reports an input price of approximately $0.10 per 1 million input tokens. However, the supplied research notes that Cohere's current public pricing page does not display an exact self-serve price for this legacy model. The reported amount should therefore be confirmed against the applicable Cohere account, contract, or cloud deployment before budgeting or production use. There is no separate output-token charge in the supplied data because the model produces embeddings rather than generated text.

Cohere provides both the standard Embed endpoint and Embed Jobs for asynchronous batch processing. Embed Jobs can be useful when a large document collection must be indexed without requiring every request to complete interactively. Exact throughput, batch limits, and deployment-specific pricing depend on the relevant endpoint or platform configuration and are not established here.

Reasoning, coding, and tool support

Embed-multilingual-v3.0 does not reason through problems or generate code. Its role is to encode content into vectors. A classification pipeline may use those vectors as features, and a retrieval system may use them to locate relevant information, but any reasoning or code generation would come from another component.

It also does not provide built-in browsing, tool calling, function calling, or code execution. This is expected for an embedding model and should be considered when designing an application architecture. A common arrangement is to use this model for retrieval and pair it with a separate generative model for response writing or agent behavior.

When to choose embed-multilingual-v3.0

Choose this model when the central requirement is multilingual semantic matching and the 512-token limit is workable. It is a reasonable fit when you need:

  • One embedding model for search across more than 100 languages.
  • Cross-lingual retrieval between queries and documents.
  • Fixed 1,024-dimensional vectors for a predictable indexing design.
  • Task-specific modes for search, classification, or clustering.
  • Image and text similarity workflows with the documented image restrictions.
  • Batch indexing through Embed Jobs.

Consider another option when you need long-context embedding, configurable vector dimensions, multiple images in a request, or richer mixed-modality processing. Cohere's embed-v4.0 is the directly relevant newer sibling identified in the research and may be a better starting point for those requirements. Choose a generative language model instead when the application must produce prose, answer questions directly, write code, use tools, or support a conversational interface.

Availability and current position

The canonical API model ID is embed-multilingual-v3.0. It is available through Cohere's Embed and Embed Jobs endpoints and is also listed for deployment through selected cloud platforms, including Amazon Bedrock. The supplied lifecycle research does not list the model as retired, although its pricing and availability should still be verified for the particular Cohere account or cloud marketplace.

Within Cohere's current model lineup, this model occupies a focused multilingual embedding role. It complements, rather than replaces, generative models such as Command: the embedding model finds and organizes relevant information, while a generative model can use that information to write an answer. Its strongest reason to remain in consideration is the combination of broad multilingual coverage, image embedding support, and a stable fixed-size vector output. Its clearest reasons not to choose it are the short 512-token input limit and its lack of generative capabilities.


Answers to Frequently Asked Questions

How does embed-multilingual-v3.0 compare with Cohere embed-v4.0?
Embed-multilingual-v3.0 offers multilingual embeddings, fixed 1,024-dimensional vectors, and image support, but it has a 512-token input limit. Embed-v4.0 may be more suitable when an application needs a larger context window, configurable vector dimensions, multiple images per request, or richer mixed-modality processing.
Does Cohere embed-multilingual-v3.0 generate answers or code?
No. It only produces embedding vectors and does not generate prose, summarize documents, write code, browse the web, call tools, or conduct conversations. A separate generative model is needed to create answers from retrieved content.
What are the input limits and supported modes of embed-multilingual-v3.0?
The maximum text input length is 512 tokens, so longer documents must be split into passages. Text supports the search_document, search_query, classification, and clustering input types. Embed v3 models also support one PNG, JPEG, WebP, or GIF image per request, supplied as a base64 data URL up to 5 MB.
What is Cohere embed-multilingual-v3.0 used for?
Cohere embed-multilingual-v3.0 converts text and supported images into 1,024-dimensional embedding vectors for semantic search, cross-lingual retrieval, classification, clustering, recommendations, RAG indexing, and image-text similarity workflows.
Can embed-multilingual-v3.0 search documents written in a different language from the user's query?
Yes. The model supports more than 100 languages and represents queries and documents in a shared vector space, allowing a query in one language to retrieve semantically similar content written in another language.


Sources 6
Provider

About Cohere