Amazon Nova Multimodal Embeddings

Amazon Nova Multimodal Embeddings

by Amazon · Active; generally available through Amazon Bedrock

Amazon Nova Multimodal Embeddings is an Amazon Bedrock model that converts text, documents, images, video, and audio into vectors in a shared semantic space. It supports cross-modal search, multimodal RAG, recommendations, classification, clustering, configurable vector dimensions, and synchronous or asynchronous processing, but does not generate conversational responses.

Embeddings Reasoning Coding
Amazon Nova Multimodal Embeddings is a specialized model available through Amazon Bedrock. It converts different kinds of content—including text, document images, photographs, video, and audio—into numerical vectors that can be compared for semantic similarity. Because the modalities share one semantic space, an application can use a text query to find relevant images, videos, documents, or audio, or compare media with other media without maintaining an entirely separate embedding workflow for each type.
Outputs

What Amazon Nova Multimodal Embeddings can produce

Embeddings
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Amazon Nova Multimodal Embeddings
Model type Other
Context window 8K tokens
Knowledge cutoff 2025-10-28
Release date 2025-10-28
Status Active; generally available through Amazon Bedrock
Knowledge cutoff notes

Amazon's AI service card states that the model was trained on data up to its release date of October 28, 2025. This is the documented training-data cutoff and is not changed by retrieval, search, or customer-provided content.

Model notes

The canonical Bedrock model ID is amazon.nova-2-multimodal-embeddings-v1:0. The model accepts text, document images, images, video, and audio and maps them into a shared semantic space. It supports embedding dimensions of 256, 384, 1024, and 3072, with 3072 as the default. Text inputs support up to approximately 8172 tokens; video and audio inputs support up to 30 seconds per input, while asynchronous processing can segment larger files. The model supports synchronous and asynchronous APIs, including provider-managed segmentation for larger text, video, and audio inputs. AWS describes pricing by modality and processing mode rather than conventional output-token pricing. Editorial scores reflect that this is a specialized embedding model and are not directly comparable to generative language models.

Cost

Model pricing

Input Modality-dependent pricing; AWS charges based on processed input and processing mode. Current pricing should be checked on the Amazon Bedrock pricing page.
Output Not applicable as output-token pricing; the model returns embeddings and pricing is based primarily on input modality and processing mode.
Model guide

Amazon Nova Multimodal Embeddings for Cross-Modal Search and RAG

Amazon Nova Multimodal Embeddings is an Amazon Bedrock embedding model that maps text, documents, images, video, and audio into a shared vector space. It is designed for cross-modal search, multimodal retrieval-augmented generation, recommendations, classification, clustering, and digital asset discovery rather than conversational text generation.

What Amazon Nova Multimodal Embeddings is

Amazon Nova Multimodal Embeddings is an embedding model provided by Amazon Web Services through Amazon Bedrock. An embedding is a numerical representation of content. Items with related meaning are placed closer together in a vector database or similarity-search system, even when they do not contain the same words or use the same media format.

The important distinction is that this model does not generate a natural-language answer. It returns vectors that other software can use for retrieval, ranking, classification, recommendation, or clustering. For example, a product team could embed a text query such as “red hiking backpack” and search a catalog of product images, while a media organization could search video and audio collections using natural-language descriptions.

Amazon positions the model for agentic retrieval-augmented generation (RAG), semantic search, recommendations, digital asset management, document classification, and clustering. In a RAG system, the model can help retrieve relevant material before a separate generative model produces an answer.

How it fits in Amazon Bedrock

The model is part of Amazon's Nova family and is accessed as a managed model in Amazon Bedrock. The documented model ID is amazon.nova-2-multimodal-embeddings-v1:0. Bedrock handles model access and provides synchronous and asynchronous invocation options, while the application remains responsible for storing vectors and performing similarity search.

This positioning makes Nova Multimodal Embeddings different from a conversational Nova model or a general-purpose language model. It is a foundation component for retrieval and organization systems, not a chatbot endpoint. A typical architecture may combine it with object storage, a vector database, application-specific ranking logic, and a separate text-generation model.

Supported content and vector output

Verified supported inputs include text, standard images, document images, video, and audio. The model maps these inputs into a common semantic space, enabling both conventional same-modality searches and cross-modal searches.

  • Text: Text inputs support up to approximately 8,172 tokens.
  • Images and document images: Useful for visual search and retrieval across scanned or image-based documents.
  • Video: Individual inputs can contain up to 30 seconds of video in the documented direct-processing workflow.
  • Audio: Individual inputs can contain up to 30 seconds of audio in the documented direct-processing workflow.

For longer text, video, and audio inputs, asynchronous processing can use provider-managed segmentation. This is more appropriate for batch ingestion or larger files than trying to process every asset as a small synchronous request.

The model supports embedding dimensions of 256, 384, 1,024, and 3,072, with 3,072 documented as the default. A larger vector can preserve more representational detail, but it also requires more storage and can increase the cost or resource requirements of vector indexing and similarity search. Smaller dimensions can be useful when storage efficiency, index size, or search speed is more important than maximum retrieval quality.

API processing and configuration

Amazon Nova Multimodal Embeddings supports synchronous invocation for smaller or latency-sensitive inputs. Its asynchronous API is intended for larger files and provider-managed segmentation, with results written to Amazon S3. The choice between the two modes is therefore operational as well as technical: synchronous requests are suitable for interactive ingestion or lookup workflows, while asynchronous processing better fits bulk media pipelines.

Requests can specify an embedding purpose, including generic indexing, retrieval, classification, or clustering. Choosing a purpose that matches the downstream task helps make the intended use explicit and allows the application to organize its embedding workflow around the job it needs to perform.

For video that contains audio, applications can request a combined audio-video embedding or separate embeddings for the audio and video components. Separate representations may be more useful when an application needs to search spoken content independently from visual scenes, while a combined representation can support queries that depend on the overall media experience.

Pricing and availability

The model became generally available on October 28, 2025, through Amazon Bedrock. It is active according to the supplied model information, and the model card states that its end-of-life date is not earlier than October 28, 2026.

A single fixed price is not supplied in the available research. AWS pricing is modality-dependent and is based on processed input and processing mode rather than conventional output-token pricing. Current AWS pricing examples identify separate modality-based charges, including per-second pricing for video processing and lower batch pricing for suitable asynchronous workloads. Because rates can vary by modality, batch or synchronous processing, and AWS Region, deployment estimates should be made using the current Amazon Bedrock pricing page rather than a generic per-request assumption.

There is no conventional maximum output-token limit because the model does not produce text. Its output is an embedding vector, and the relevant output choices are the supported dimensions: 256, 384, 1,024, or 3,072.

Main strengths

  • Cross-modal retrieval: Text, images, documents, video, and audio can be represented in one semantic space, reducing the need to coordinate unrelated embedding models for basic cross-media search.
  • Broad media coverage: The model supports content types that are often split across separate search systems, including document images, video, and audio.
  • Flexible vector sizes: Four embedding dimensions allow teams to balance retrieval quality, storage requirements, and search infrastructure costs.
  • Batch-friendly processing: Asynchronous invocation and segmentation support larger files and high-volume ingestion workflows.
  • Task-specific configuration: Embedding purposes cover indexing, retrieval, classification, and clustering use cases.

These strengths are especially relevant when an organization has a mixed media library. A single search experience can potentially retrieve a scanned page, a product photograph, a video segment, or an audio recording from a natural-language query.

Limitations and trade-offs

Nova Multimodal Embeddings is not a conversational or generative model. It does not write explanations, answer questions, generate code, create images, synthesize speech, or return confidence scores for its vectors. An application that needs an answer in prose must add a separate generation or question-answering component.

Embedding quality is also dependent on the data and the retrieval task. AWS recommends evaluating the model with representative customer data because background conditions, lighting, camera angle, audio quality, speaker accent, professional terminology, and other characteristics can affect similarity results. A visually noisy image collection or highly specialized vocabulary may require careful testing and preprocessing.

The model does not replace a vector database or another similarity-search system. Teams must select an index, store vectors and metadata, manage ingestion, and decide how to filter or rank results. Larger vectors may improve retrieval in some workloads but increase storage and indexing demands. Smaller vectors can reduce infrastructure costs, but their suitability should be validated against representative search queries.

Direct processing limits also matter. Text supports approximately 8,172 tokens, while individual video and audio inputs support up to 30 seconds in the documented workflow. Longer content requires segmentation or asynchronous processing, which adds pipeline complexity and may not be appropriate for an interactive request that needs an immediate result.

Reasoning, coding, and tool support

This model has no conversational reasoning capability in the usual language-model sense. Its role is to calculate semantic representations, not to carry out multi-step reasoning or explain why two items are similar. The supplied editorial assessment rates reasoning and coding capability at 1 out of 10 because those capabilities are outside the model's intended function, not because the model is a poor choice for its embedding purpose.

It does not provide built-in tool or function calling, streaming responses, or generative JSON output. It also does not support fine-tuning or caching according to the supplied model information. Batch-style processing is supported through the asynchronous workflow, but that should not be confused with a conversational batch-generation API.

Best use cases

  • Cross-modal media search: Find photographs, video, or audio from text descriptions, or locate visually and semantically related assets.
  • Multimodal RAG: Retrieve relevant pages, diagrams, images, and media before passing the results to a separate answer-generation model.
  • Digital asset management: Organize and discover large collections of marketing images, recordings, documents, and video.
  • Commerce recommendations: Match product descriptions, customer queries, catalog images, and related items in a shared search space.
  • Classification and clustering: Group mixed-media content or assign content to categories based on semantic similarity.

When to choose this model

Choose Amazon Nova Multimodal Embeddings when the central problem is finding, organizing, or comparing mixed media and when cross-modal retrieval is valuable. It is a strong fit for teams already using Amazon Bedrock and AWS storage or data services, particularly when asynchronous ingestion and configurable vector dimensions are useful.

Choose a text-only embedding model instead when the corpus and queries are exclusively text and there is no need to search images, audio, video, or document visuals. A text-focused option may provide a simpler and potentially more economical pipeline for that narrower task, although the supplied research does not establish a specific competing model or price.

Choose a generative language or multimodal chat model when the application must answer questions, summarize content, generate code, explain results, or produce natural-language responses. Nova Multimodal Embeddings can supply the retrieval layer for such an application, but it is not a substitute for the generation layer.

Practical evaluation checklist

Before deployment, test the model with real queries and representative media rather than relying only on a generic similarity score. Compare retrieval quality at the available vector dimensions, measure the effect of synchronous versus asynchronous processing, and account for the storage and indexing impact of the selected vector size. For video and audio, test whether combined embeddings or separate modality-specific embeddings better match the search behavior users expect.

Also estimate costs by modality and processing mode. Video and audio workloads can behave differently from text and image workloads, and longer files may require segmentation. Finally, treat retrieved results as evidence for a downstream system rather than as explanations produced by the embedding model itself. The model's value is in structuring content for efficient semantic comparison.


Answers to Frequently Asked Questions

When should organizations use Amazon Nova Multimodal Embeddings?
Organizations should consider it when they need to search, compare, classify, or organize mixed media such as text, images, documents, video, and audio. It is particularly suitable for cross-modal search, multimodal RAG, digital asset management, commerce recommendations, and large-scale media clustering.
Is Amazon Nova Multimodal Embeddings a generative or conversational AI model?
No. The model generates embedding vectors rather than natural-language answers, explanations, code, or other content. Applications that need conversational responses must combine it with a separate generative model, typically using the embeddings for retrieval in a RAG pipeline.
What are the supported vector dimensions for Amazon Nova Multimodal Embeddings?
Amazon Nova Multimodal Embeddings supports vector dimensions of 256, 384, 1,024, and 3,072, with 3,072 documented as the default. Smaller vectors can reduce storage and search costs, while larger vectors may preserve more representational detail.
What types of content does Amazon Nova Multimodal Embeddings support?
The model supports text, standard images, document images, video, and audio. It can represent these formats in a shared semantic space, enabling searches such as matching a text query with product images or finding video and audio content from natural-language descriptions.
What is Amazon Nova Multimodal Embeddings?
Amazon Nova Multimodal Embeddings is an AWS embedding model available through Amazon Bedrock. It converts text, images, document images, video, and audio into numerical vectors that can be used for semantic search, retrieval-augmented generation, recommendations, classification, and clustering.


Sources 6
Provider

About Amazon