Cohere Embed v4.0 is a multimodal embedding model from Cohere, released on April 15, 2025, and listed as generally available. Instead of generating an answer, it converts supplied content into numerical vector representations called embeddings. Those vectors let software compare meaning, helping a search system find relevant passages, images, document pages, charts, or tables even when the wording is not an exact match.
The model is aimed primarily at enterprise retrieval and indexing. It can process text, images, and mixed text-image inputs, including image-rich documents such as PDFs. This makes it a fit for semantic search, retrieval-augmented generation (RAG), document discovery, classification, clustering, and image-to-text retrieval. It is not a chatbot or a general-purpose generative model.
What is Cohere Embed v4.0?
An embedding model maps input content into a vector: a list of numbers that captures useful relationships in the content. A search application can embed a user's query and compare it with stored document embeddings. Content with nearby vectors is treated as semantically related, even if it uses different words.
Embed v4.0 applies this approach across more than plain text. A business could index written reports, scanned pages, presentation slides, figures, tables, and charts in a common retrieval workflow. The model can also represent mixed inputs containing text and images, which is useful when the meaning of a document depends on both its prose and its visual material.
Embed v4.0 belongs to Cohere's Embed model family, alongside the provider's broader Command, Rerank, Parse, and Aya offerings. Its specific role is representation and retrieval: it supplies vectors that other search, ranking, classification, or generative systems can use. It does not produce a textual response, image, audio file, or video.
Supported inputs and outputs
The verified input modalities are text and images. Mixed text-and-image content is also supported. Cohere's documentation describes use cases involving document screenshots, figures, tables, charts, and PDF pages. Audio and video inputs are not listed for this model.
The output is an embedding vector rather than generated content. Embed v4.0 supports four output dimensions: 256, 512, 1,024, and 1,536, with 1,536 as the default. Lower dimensions can reduce storage and similarity-search costs, while larger vectors retain a more detailed representation. This configurable-dimension approach is commonly called Matryoshka representation learning, although the important practical point is that the same model can be adapted to different indexing budgets.
Supported embedding formats include float, int8, uint8, binary, and ubinary. Quantized or binary formats can reduce storage and accelerate some vector-search workloads, but the best choice depends on the database and the accuracy requirements of the application.
Context and request limits
Embed v4.0 has a 128,000-token context length. This is a substantial input window for long documents and makes it possible to handle large content units without automatically splitting every document into very small pieces. In practice, document segmentation may still be useful when search needs precise passage-level results or when a downstream system has its own size limits.
The Embed API supports up to 96 text or mixed-content inputs in one call. For Embed v4.0 image requests, the combined image payload limit is 20 MB. Cohere's documentation indicates that Embed v4.0 does not have the v3-style one-image-per-call restriction, so an image-bearing request can contain multiple supported image inputs within the documented request limits.
There is no maximum generated-output-token figure because the model does not generate text. It returns embeddings, and the output size is determined by the selected vector dimension and format.
Pricing and deployment
Cohere's listed usage pricing is based on embedded input tokens. The supplied pricing information lists $0.12 per 1 million text input tokens and $0.47 per 1 million image tokens. There is no generative output-token charge for Embed v4.0 because it produces vectors rather than generated language.
These prices apply to API-style embedding usage and should not be confused with dedicated deployment pricing. Cohere also lists Model Vault deployment tiers for Embed 4: Small at $4 per hour or $2,500 per month, and Medium at $5 per hour or $3,250 per month. The supplied information does not establish whether those monthly figures represent a month-to-month commitment or an annual-equivalent rate, so organizations should confirm the commercial terms with Cohere.
Embed v4.0 is available through Cohere Platform and, according to the release information, through AWS SageMaker and Azure AI Foundry. Cohere also supports enterprise deployment approaches such as private infrastructure and Model Vault, which may be relevant when data isolation, governance, or infrastructure control matters more than the lowest per-token price.
What Embed v4.0 does well
- Multimodal retrieval: It can place text, images, and mixed-content documents into retrieval workflows instead of limiting indexing to written passages.
- Long-document handling: The 128,000-token context length supports large inputs and document pages with substantial surrounding context.
- Enterprise search: Its design fits internal knowledge bases, document repositories, PDF search, and RAG pipelines where results must be retrieved before a separate model generates an answer.
- Flexible vector sizing: Four dimensions let teams balance retrieval quality, storage, and search cost.
- Multilingual use: Cohere positions the model for multilingual embeddings, making it suitable for search systems that span multiple languages.
- Operational flexibility: Multiple formats, batch-style input limits, and deployment options can help teams adapt the model to an existing vector infrastructure.
These are capability and positioning observations based on the supplied documentation. They are not claims that Embed v4.0 wins every independent benchmark. The research does not provide benchmark scores, and the model's editorial reasoning and coding scores are not meaningful measures of generative reasoning or programming ability.
Reasoning, coding, and tool support
Embed v4.0 is not a reasoning model in the conversational sense. It does not plan through a problem, explain a solution, or decide which tool to call. Its job is to encode supplied content so another component can retrieve or compare it.
It can embed source code as text if an application needs to search a code repository, but it is not a code-generation model and should not be selected to write, debug, or refactor software. The supplied specifications list no tool or function-calling support, no structured-output mode, and no generative JSON capability. A typical architecture would pair Embed v4.0 with a vector database, retrieval logic, and a separate language model when a natural-language answer is required.
Best use cases
Enterprise document search
Embed v4.0 is well suited to indexing internal policies, reports, presentations, manuals, and other business records. Its multimodal inputs are especially useful when a document's important information appears in tables, diagrams, or screenshots rather than in surrounding text alone.
Multimodal retrieval-augmented generation
In a RAG system, Embed v4.0 can retrieve relevant text passages or visual document elements for a separate generative model. For example, a question about a chart in a financial report could retrieve the page image or related mixed-content representation before an answer is composed.
Classification and clustering
Embedding vectors can support grouping documents by similarity or assigning them to categories. This can help organize large repositories, discover duplicate or related material, and route incoming content to downstream workflows.
Image-to-text and cross-modal retrieval
Because the model handles text and images in a shared retrieval workflow, an application can search visually rich material using textual queries, subject to the accuracy and indexing design of the complete system.
Limitations and trade-offs
The most important limitation is that Embed v4.0 does not answer questions directly. A search result is not a finished response, so applications that need explanations, summaries, code, or conversational interaction must add another model and orchestration layer.
Multimodal support can also increase operational complexity. Teams need to decide how to represent PDF pages, whether to preserve page images alongside extracted text, which vector dimension to use, and whether compressed formats provide an acceptable quality-cost balance. The 20 MB combined image limit and 96-input request limit must be considered when designing ingestion jobs.
Embed v4.0 may be a poor choice for workloads that only need simple keyword matching, direct text generation, image creation, speech, transcription, or tool-using agents. A conventional text embedding model may be sufficient for a text-only corpus, while a generative model is more appropriate when the primary requirement is to produce an answer. The supplied research does not identify a specific competing model that should replace Embed v4.0 in every case.
When to choose Cohere Embed v4.0
Choose Embed v4.0 when a search or indexing system needs one embedding model for multilingual text, images, and mixed-content documents; when long inputs are important; or when vector dimensions and formats need to be adjusted for storage and retrieval constraints. It is particularly compelling for enterprise repositories where PDFs, charts, presentations, and scanned or visual material are part of the information users need to find.
Consider another option when the application needs generated text rather than vectors, direct reasoning, code generation, speech, video, or image creation. For a purely text-based collection, a text-only embedding model may offer a simpler or cheaper design. For an interactive assistant, Embed v4.0 should generally be treated as the retrieval component rather than the complete assistant.
Bottom line
Cohere Embed v4.0 is a specialized, generally available embedding model for multimodal enterprise retrieval. Its defining features are support for text, images, and mixed-content inputs; a 128,000-token context length; configurable 256-to-1,536-dimensional vectors; multiple embedding formats; and pricing based on embedded input tokens. It is a strong fit for semantic search, multimodal RAG, document indexing, classification, and clustering, but it should not be evaluated as a chatbot or generative language model.

