Amazon Titan

Amazon Titan Multimodal Embeddings G1

by Amazon · Active

Amazon Titan Multimodal Embeddings G1 converts text, images, or combined text and image inputs into shared vector representations. It supports 256-, 384-, and 1,024-dimensional embeddings for text-to-image search, image similarity, recommendations, personalization, and specialized image-text matching through Amazon Bedrock.

Embeddings Reasoning Coding
Amazon Titan Multimodal Embeddings G1 gives applications one vector representation for text and images. A product image, a natural-language description, or both together can be transformed into an embedding that is compared with other stored embeddings. This makes the model useful for text-to-image retrieval, visual catalog search, recommendations, and image similarity workflows through Amazon Bedrock.
Outputs

What Amazon Titan Multimodal Embeddings G1 can produce

Embeddings
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Amazon Titan
Model type Multimodal
Context window 256 tokens
Release date 2023-11-29
Status Active
Knowledge cutoff notes

AWS does not publish a directly verifiable knowledge-cutoff date for this embedding model in the cited model documentation. Embedding models are not conversational knowledge models, and the documented input limits and training/customization details should not be interpreted as a knowledge cutoff.

Model notes

Canonical Amazon Bedrock model ID is amazon.titan-embed-image-v1. The model accepts inputText, inputImage, or both. Combined text-and-image requests produce an embedding based on the average of the text and image embeddings. Inference supports 256, 384, or 1,024 output dimensions, with 1,024 as the default. AWS documents a maximum inference image size of 25 MB and 2,048 × 2,048 pixels, a maximum text input of 256 tokens, and English language support. Bedrock customization uses image-text pairs and supports PNG and JPEG training images. The model card lists the lifecycle as Active and gives an EOL date no sooner than November 29, 2024, without specifying a definitive shutdown date.

Cost

Model pricing

Input Text: $0.0008 per 1,000 input tokens; images: $0.00006 per input image. Pricing may vary by AWS Region and current Bedrock pricing terms.
Output No separate output-token price; the model returns embedding vectors.
Model guide

Amazon Titan Multimodal Embeddings G1 for Text-to-Image Search and Similarity

Amazon Titan Multimodal Embeddings G1 is an Amazon Bedrock embedding model that converts text, images, or combined text-and-image inputs into vectors in a shared semantic space. Its main use is connecting text and visual content for multimodal search, recommendations, personalization, and similarity matching rather than generating natural-language responses or media.

What Amazon Titan Multimodal Embeddings G1 does

Amazon Titan Multimodal Embeddings G1 is an embedding model provided by Amazon Web Services through Amazon Bedrock. Instead of generating an answer, an image, or another piece of content, it converts input into a numerical vector. A vector is a list of numbers that represents the meaning or visual characteristics of the input in a form that software can compare.

The model's distinguishing feature is that text and images are mapped into a shared semantic space. An application can therefore compare a text description with image vectors, compare one image with another, or combine text and image information in a single retrieval workflow. For example, a retailer could index product photographs and allow shoppers to search with phrases such as “black leather boots,” upload a reference image, or use both an image and a description.

The model is identified in Bedrock by the canonical model ID amazon.titan-embed-image-v1. It belongs to Amazon's Titan model family, but it is specifically an embedding model rather than a conversational or generative model.

Where it fits in Amazon's model lineup

Titan Multimodal Embeddings G1 is designed for applications that need retrieval-oriented representations of visual and textual content. It is separate from Amazon's text-generation, conversational, and image-generation offerings. Its output is intended to be stored in a vector index and used by a search, recommendation, or matching system.

This positioning matters when selecting a model. A generative model may be more appropriate for writing product descriptions, answering questions, or analyzing a document in natural language. Titan Multimodal Embeddings G1 is the more relevant choice when the central task is finding items that are semantically or visually similar to a query.

Supported inputs and output

The model accepts at least one of two input types: inputText or inputImage. It supports:

  • Text-only requests, which produce text embeddings.
  • Image-only requests, which produce image embeddings.
  • Combined text-and-image requests, which produce an embedding based on the average of the text and image embeddings.

Images are supplied as base64-encoded data in the Bedrock runtime request. The output is a floating-point embedding vector. The model does not directly return natural-language text, generated images, audio, or video, so an application needs separate retrieval, ranking, or generation components if it must present a user-facing response.

SpecificationVerified detail
ProviderAmazon Web Services
Bedrock model IDamazon.titan-embed-image-v1
Input modalitiesText, images, or text plus images
OutputFloating-point embedding vector
Embedding dimensions256, 384, or 1,024
Default dimensions1,024
Supported languageEnglish
Inference optionsOn-Demand and Provisioned Throughput

Input limits and technical constraints

The maximum text input is 256 tokens. This is a relatively short limit for document-oriented embedding, so long documents generally need to be divided into smaller passages or summarized before they are embedded. The supplied model documentation does not describe Titan Multimodal Embeddings G1 as a long-context document embedding model.

For inference, an image can be up to 25 MB and up to 2,048 × 2,048 pixels. Supported image workflows include image-to-image similarity and text-to-image retrieval, provided the application handles image encoding, vector storage, indexing, and similarity search.

The choice of vector dimension affects storage and search configuration. A 1,024-dimensional vector contains more values than a 256-dimensional vector and may require more storage and computational resources. Smaller dimensions can be useful when reducing index size or search overhead is a priority, but the supplied research does not establish a benchmark showing how accuracy changes between the available sizes. In practice, the selected dimension should be chosen alongside the vector database configuration, and changing dimensions later may require a separate index or a re-embedding workflow.

What it is useful for

One of the clearest use cases is searching a visual catalog with language. A stock-photo library, media archive, product catalog, or enterprise asset repository can embed its images and then compare them with a user's text query. A combined text-and-image query can add constraints that are difficult to express through either modality alone.

For example, a user might submit a photo of a chair and add “in a lighter fabric.” The application can use the combined representation to retrieve visually related catalog items while incorporating the text description.

Image similarity and duplicate-style discovery

Image-only embeddings allow applications to find visually related content. This can support reference-image search, related-product discovery, asset organization, and visual browsing. The model supplies the representation, while the surrounding application determines how similarity is measured, which database is used, and how results are ranked.

Recommendations and personalization

Shared text-and-image vectors can help match users with products or content when catalog items have both descriptive text and imagery. A recommendation system might compare an item a user viewed with other catalog vectors, or combine textual product information with visual characteristics when selecting related products.

Domain-specific visual and textual matching

Amazon Bedrock supports customization using image-text pairs. This is relevant when the base model does not adequately represent an organization's specialized visual vocabulary, product catalog, or image-text relationships. The supplied documentation describes training datasets from 1,000 to 500,000 examples and validation datasets from 8 to 50,000 examples. Training images can use PNG or JPEG format and have a 25 MB image-size limit.

Customization considerations

Customization is not the same as turning the model into a general-purpose language or vision assistant. It is intended to adjust how the embedding space represents a particular domain. An organization with specialized products, industry terminology, or recurring visual patterns may use image-text pairs to make similarity and retrieval more relevant to its own data.

Customization also introduces operational work. Teams need suitable paired examples, a validation set, a process for creating and updating vectors, and an index configured for the selected output dimension. The supplied research confirms that Bedrock supports this customization path, but it does not provide a benchmark or guaranteed improvement for a particular dataset.

Pricing and access

AWS documents usage-based pricing for Titan Multimodal Embeddings G1. The supplied pricing information lists $0.0008 per 1,000 text input tokens and $0.00006 per input image. Pricing can vary by AWS Region and may change under current Amazon Bedrock pricing terms.

There is no separate output-token price in the supplied research because the model returns embedding vectors rather than generated text. The total cost of a production system can still include vector database storage, indexing, retrieval, application infrastructure, and any additional Bedrock services used around the embedding workflow.

Access is through the Amazon Bedrock runtime, with availability dependent on AWS Region and account configuration. Developers should verify that the model is enabled in the intended deployment region before designing a production system around it.

Capabilities, speed, and trade-offs

This model is not a reasoning model in the conversational sense. It does not carry out multi-step explanations, answer questions, or use external tools. Its purpose is to produce representations that other software can compare. The supplied evaluation records a reasoning score of 1 and a coding score of 1; these are catalog evaluations, not provider-published benchmark results, and should not be interpreted as measures of embedding quality.

The catalog also records a speed score of 8 and a cost score of 8. These are editorial or database-level assessments rather than AWS guarantees. The practical speed and cost of a deployment depend on request volume, image size, selected vector dimension, AWS Region, inference option, and the surrounding search infrastructure.

Compared with a generative model, Titan Multimodal Embeddings G1 can be a more direct and economical fit when the task is repeated vector creation for search or matching. Compared with a text-only embedding model, it is better suited when images are central to the query or catalog. However, its English focus and 256-token text limit make a dedicated long-document or multilingual embedding option more appropriate for large multilingual knowledge bases.

Important limitations

  • It does not generate answers: the output is an embedding vector, so applications need a separate retrieval and presentation layer.
  • English support: the documented language support is English; the supplied research does not establish broad multilingual capability.
  • Short text limit: the maximum input is 256 tokens, which is restrictive for whole-document embedding.
  • Image handling is application work: images must be encoded and supplied within the documented size limits, while storage, indexing, similarity calculations, and result ranking are managed by the application.
  • Dimension changes affect infrastructure: switching between 256, 384, and 1,024 dimensions may require a separate vector index or re-embedding process.
  • No native tool or function use: the model does not perform actions, browse the web, call functions, or stream conversational output according to the supplied specifications.
  • No direct media generation: it does not produce text, images, audio, or video as model output.

When to choose Amazon Titan Multimodal Embeddings G1

Choose Titan Multimodal Embeddings G1 when an application needs text and images to participate in the same retrieval or similarity system. It is a strong fit for visual product catalogs, text-to-image search, image similarity, multimodal recommendations, personalization, and specialized image-text matching on AWS.

It is less suitable when the primary need is conversational reasoning, code generation, document summarization, multilingual semantic search, or content generation. In those cases, a generative model, a long-context text embedding model, or a multilingual embedding model may be more appropriate. The key selection question is whether the application needs a shared vector space for visual and textual inputs. If it does, the model's Bedrock integration, configurable dimensions, and image-text customization support are its main practical advantages.

Bottom line

Amazon Titan Multimodal Embeddings G1 is a focused embedding model for connecting text and images in search and recommendation systems. Its 256-, 384-, and 1,024-dimensional outputs, support for combined inputs, and Bedrock customization path make it useful for multimodal retrieval. Its limitations are equally important: it is English-focused, accepts only 256 text tokens, returns vectors rather than answers, and requires the application to provide the surrounding vector-search infrastructure.


Answers to Frequently Asked Questions

What is Amazon Titan Multimodal Embeddings G1 used for?
Amazon Titan Multimodal Embeddings G1 converts text, images, or combined text and image inputs into numerical embedding vectors. These vectors can power text-to-image search, image similarity, recommendations, personalization, and multimodal retrieval systems.
What are the input and output limits of Titan Multimodal Embeddings G1?
The model accepts text, images, or both. Text input is limited to 256 tokens, while images can be up to 25 MB and 2,048 × 2,048 pixels. It returns a floating-point embedding with 256, 384, or 1,024 dimensions, with 1,024 as the default.
Can Amazon Titan Multimodal Embeddings G1 search images using text?
Yes. The model maps text and images into a shared semantic space, allowing applications to compare a text query with image embeddings for tasks such as product, stock-photo, and media-asset search.
What is the Bedrock model ID for Titan Multimodal Embeddings G1?
The canonical Amazon Bedrock model ID is amazon.titan-embed-image-v1.
What are the main limitations of Amazon Titan Multimodal Embeddings G1?
The model is documented for English, has a 256-token text limit, and returns vectors rather than generated answers or media. Applications must provide image encoding, vector storage, indexing, similarity measurement, ranking, and any user-facing response generation.


Sources 6
Provider

About Amazon