Embed v3.0

embed-english-v3.0

by Cohere · Live

An English-focused Cohere embedding model that creates 1,024-dimensional vectors from text and supported images for semantic search, retrieval, classification, clustering, recommendations, and similarity matching.

Embeddings
Cohere Embed English v3.0 is a specialized representation model rather than a chatbot or text-generation system. It converts English text, and supported images, into numerical vectors that applications can compare to find semantic relationships. This makes it useful for search, retrieval-augmented generation, classification, clustering, recommendations, and duplicate detection. The model is available through Cohere's Embed API and Embed Jobs endpoint, and its v3.0 family supports image embeddings in addition to English text.
Outputs

What embed-english-v3.0 can produce

Embeddings
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Embed v3.0
Model type Other
Context window 512 tokens
Release date 2023-11-02
Status Live
Knowledge cutoff notes

Cohere does not publish a model knowledge-cutoff date for Embed English v3.0. As an embedding model, its documented behavior is based on vector representation rather than conversational knowledge responses.

Model notes

Cohere's current model documentation lists 1,024-dimensional embeddings, a 512-token context length, English text support, image input, cosine similarity, dot-product similarity, Euclidean distance, and Embed plus Embed Jobs endpoints. The v3.0 family gained image-embedding support on October 22, 2024. Image embedding accepts PNG, JPEG, WebP, or GIF data up to 5 MB as a base64 data URL and does not support batching for v3.0. Cohere's current public pricing page does not prominently display a dedicated per-token price for this model; $0.10 per million input tokens is the commonly reported production rate. The model is active in Cohere's current catalog and status page. It has no generative output, so max output tokens and output-token pricing are not applicable.

Cost

Model pricing

Input $0.10 per 1 million input tokens
Output Not applicable; the model returns embeddings rather than billed generated text
Model guide

Cohere Embed English v3.0 for English Semantic Search and Image Embeddings

Cohere Embed English v3.0 is an English-focused embedding model for semantic search, retrieval, classification, clustering, recommendations, and similarity matching. It produces 1,024-dimensional vectors from English text and supported individual image inputs, with a 512-token context limit and no generative output.

What Cohere Embed English v3.0 does

Cohere Embed English v3.0 converts supported input into an embedding: a list of numbers representing the meaning or features of that input. The resulting vectors can be stored in a vector database and compared with other vectors using cosine similarity, dot-product similarity, or Euclidean distance.

In practical terms, an application can embed a user's search query and compare it with embedded documents to find relevant passages, even when the wording is different. For example, a query about “reducing cloud infrastructure costs” can match a document that discusses “lowering hosting expenses” because the model is designed to represent related meaning rather than only identical keywords.

The model is provided by Cohere and belongs to the company's Embed v3.0 family. It is a hosted, closed model intended for use through Cohere's services and supported deployment integrations. It is not a conversational model and does not return prose, code, generated images, audio, or video.

Verified specifications and supported inputs

SpecificationDetails
Model IDembed-english-v3.0
ProviderCohere
Model typeEnglish embedding and representation model
Release dateNovember 2, 2023
Vector size1,024 dimensions
Context limit512 tokens
Text inputEnglish text
Image inputSupported images through the v3.0 multimodal update
Generated outputNone; the model returns embeddings
Primary accessCohere Embed API and Embed Jobs endpoint

The 512-token context limit is important when embedding long documents. Content longer than the model's supported input window must be divided into smaller chunks before indexing. Chunking can improve retrieval precision, but it also means the application must preserve useful metadata and manage how retrieved passages are assembled later.

Cohere documents input-type options for different workloads, including search queries, searchable documents, classification, and clustering. These options help the model produce representations suited to the role of the input. A search system should therefore distinguish between the query being searched and the documents being indexed rather than treating every input as interchangeable.

Image embedding support

Embed English v3.0 originally focused on English text. Cohere later added image-embedding support to the v3.0 family, making it possible to represent supported images as vectors. The model still does not generate images: image input produces an embedding that can be compared with other image or text embeddings for retrieval and similarity workflows.

According to Cohere's documentation, supported image formats include PNG, JPEG, WebP, and GIF. Images can be supplied as base64-encoded data in a data URL, and the maximum image size is 5 MB. Image embedding for v3.0 does not support batching, so an image request is limited to one image at a time.

This capability can support use cases such as finding visually or semantically related images, matching image content with text descriptions, and building search systems that combine written metadata with visual inputs. The supplied research does not establish that the model performs image captioning, object detection, or image generation, so those should not be treated as supported outputs.

Main use cases

  • Semantic search: Find relevant content based on meaning instead of exact keyword overlap.
  • Document retrieval: Index passages for search, enterprise knowledge systems, or retrieval-augmented generation.
  • Vector database indexing: Store 1,024-dimensional representations for nearest-neighbor searches.
  • Classification: Use embeddings as representations for English text classification workflows.
  • Clustering: Group related documents, support tickets, queries, or other English-language content.
  • Recommendations and similarity matching: Compare products, documents, records, or images according to their represented content.
  • Duplicate detection: Identify items that are semantically similar even when their wording differs.
  • Multimodal retrieval: Compare supported image and text representations where the application's workflow is designed for that purpose.

For retrieval-augmented generation, Embed English v3.0 handles the retrieval stage rather than the answer-writing stage. An application can embed a question, retrieve nearby document chunks, and pass those chunks to a separate generative model. Embed English v3.0 itself does not reason over the retrieved material or produce the final response.

Strengths and practical trade-offs

The model's main strength is specialization. It focuses on producing consistent English representations for search and similarity tasks instead of spending resources on conversational generation. Its 1,024-dimensional vectors provide a fixed format for indexing, and its query-document input types are useful for search systems that need different representations for questions and documents.

The later image support broadens the model beyond text-only indexing. A system can use the same v3.0 family for certain text and image retrieval tasks, although image requests have their own format, size, and batching restrictions.

Embedding systems are generally faster and less expensive for retrieval than using a generative language model to inspect every document or answer every search query. However, the exact speed and cost depend on request volume, document chunking, infrastructure, and the surrounding vector-search system. The editorial assessment supplied for this model rates its speed at 7 out of 10 and cost at 8 out of 10; these are comparative editorial scores, not Cohere-published benchmarks or guarantees.

The model also has clear trade-offs. It is English-focused, limited to 512 tokens per input, and does not offer configurable output dimensions according to the supplied research. It has no text-generation output, no tool or function calling, no streaming response, and no reasoning or coding mode. Applications that need multilingual retrieval, long inputs, generated answers, or broader document handling may need a newer or different model.

Pricing and availability

Embedding usage is billed by the amount of input processed rather than by generated output tokens. Public model catalogs and third-party references commonly report approximately $0.10 per one million input tokens for Embed English v3.0. Cohere's current public pricing page does not prominently display a dedicated per-token price for this legacy model, so the reported figure should be verified with Cohere before production budgeting or procurement.

There is no output-token charge for generated text because Embed English v3.0 does not generate text. The model is available through Cohere's Embed API and Embed Jobs endpoint and is listed in Cohere's current model catalog and operational status documentation. Cohere also documents deployment integrations such as Amazon Bedrock, where the corresponding identifier is cohere.embed-english-v3.

Public API access begins with Cohere's developer access and trial arrangements, but production use is subject to account, billing, and service terms. Deployment availability and pricing can differ by platform, region, and commercial arrangement.

Capabilities it does not have

Embed English v3.0 should not be evaluated as a general-purpose AI assistant. It does not provide:

  • Chat or conversational responses
  • Text, code, image, audio, or video generation
  • Tool use or function calling
  • Web search or browsing
  • Audio or video input
  • Configurable reasoning levels
  • Streaming generated responses
  • Structured JSON output as a generative response

Its output is an embedding vector. An application must supply the surrounding components, such as a vector index, similarity function, retrieval logic, classifier, or separate generative model.

When to choose Embed English v3.0

Choose Cohere Embed English v3.0 when the primary requirement is English semantic representation and the 512-token context limit is acceptable. It is a reasonable fit for an established search or retrieval pipeline that already expects 1,024-dimensional vectors, needs Cohere's query-document input specialization, or benefits from supported image embeddings.

It may also be suitable when cost and retrieval throughput matter more than generative capabilities. Embedding a large collection once and searching the resulting vectors is usually more appropriate than repeatedly sending the entire collection to a generative model.

Consider another option when the application needs multilingual coverage, longer inputs, selectable vector dimensions, or mixed document formats. Cohere Embed v4.0 and Cohere's multilingual Embed models are relevant alternatives mentioned in the supplied research, but the right choice depends on the required languages, context size, input formats, deployment target, and current pricing.

Use a separate generative model when the application must write answers, summarize retrieved passages, generate code, or hold a conversation. Use a separate vision, audio, or video system when the task requires analysis beyond the supported image-embedding workflow.

Bottom line

Cohere Embed English v3.0 is a focused embedding model for English retrieval and similarity workloads. It produces 1,024-dimensional vectors, accepts English text and supported individual images, and supports a 512-token context window. Its value comes from efficient representation for search, classification, clustering, and retrieval pipelines rather than from reasoning or generation. It remains a practical choice for compatible English-language systems, while newer or multilingual alternatives deserve comparison when the project requires broader inputs, larger context, or configurable vector output.


Answers to Frequently Asked Questions

How much does Cohere Embed English v3.0 cost?
Public model catalogs and third-party references commonly report approximately $0.10 per one million input tokens, but Cohere's current pricing should be verified before production budgeting because pricing can vary by platform, region, account, and commercial agreement.
Can Cohere Embed English v3.0 generate answers or summaries?
No. Cohere Embed English v3.0 only produces embedding vectors and does not generate answers, summaries, code, images, audio, or video. A separate generative model is needed to write responses from retrieved content.
Does Cohere Embed English v3.0 support image embeddings?
Yes. The model supports image embeddings for PNG, JPEG, WebP, and GIF files up to 5 MB. Images can be submitted as base64-encoded data URLs, but image requests do not support batching and are limited to one image at a time.
What are the main specifications of Cohere Embed English v3.0?
The model ID is embed-english-v3.0. It produces 1,024-dimensional vectors, supports up to 512 tokens per input, accepts English text and supported images, and returns embeddings rather than generated text.
What is Cohere Embed English v3.0 used for?
Cohere Embed English v3.0 converts English text and supported images into embedding vectors for semantic search, document retrieval, vector database indexing, classification, clustering, recommendations, similarity matching, duplicate detection, and multimodal retrieval.


Sources 8
Provider

About Cohere