Tongyi Embedding Vision

tongyi-embedding-vision-plus

by Qwen · Available

Tongyi Embedding Vision Plus is Alibaba Cloud's multimodal embedding model for converting text, images, videos, and multi-image sequences into 1,152-dimensional vectors. It supports cross-modal retrieval, similarity search, classification, recommendation, and semantic indexing, with input limits and pricing documented for the international Singapore deployment.

Embeddings Reasoning Coding
Tongyi Embedding Vision Plus is a vision-focused multimodal embedding model available through Alibaba Cloud Model Studio. It is designed to place text, images, videos, and image sequences into a numerical representation that applications can compare. This makes it suitable for search and recommendation systems that need to connect different media types, such as finding product images from a text query or locating videos that match an image.
Outputs

What tongyi-embedding-vision-plus can produce

Embeddings
Inputs

What it can understand

Text Images Video Multimodal input
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Tongyi Embedding Vision
Model type Multimodal Embedding
Release date 2025-10-31
Status Available
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this embedding model.

Model notes

Tongyi Embedding Vision Plus is an independent multimodal embedding model. It accepts Chinese and English text, JPG/PNG/BMP images, videos through public URLs, and multi-image sequences of up to eight images. Text input is limited to 1,024 tokens per request; images are limited to 3 MB each and videos to 10 MB. The default and documented embedding dimension is 1,152. The model returns dense vectors and does not generate text or other media. Alibaba Cloud documentation lists no support for function calling, structured outputs, web search, context caching, batch inference, or fine-tuning for this exact model. Editorial scores are comparative estimates for an embedding model rather than vendor benchmarks.

Cost

Model pricing

Input $0.09 per 1 million input tokens for text, image, and video in the international Singapore deployment
Output $0; embedding output is free
Model guide

Tongyi Embedding Vision Plus for Cross-Modal Image and Video Retrieval

Tongyi Embedding Vision Plus is Alibaba Cloud's multimodal embedding model for turning Chinese and English text, images, videos, and multi-image sequences into 1,152-dimensional vectors. These vectors support cross-modal retrieval, semantic similarity, classification, recommendation, and multimodal indexing rather than content generation.

What is Tongyi Embedding Vision Plus?

Tongyi Embedding Vision Plus is an independent multimodal embedding model from Alibaba Cloud Model Studio. Instead of generating a written answer or creating an image, it converts supported content into dense numerical vectors called embeddings. Applications can compare those vectors to estimate semantic similarity and retrieve related items from an indexed collection.

The model is intended for workflows where different media types need to be searched together. For example, a product catalog could store image embeddings and use a text description as the search query. A media library could compare an uploaded image with video embeddings, while a recommendation system could match products, images, and descriptions according to their meaning rather than only their exact keywords.

According to the supplied documentation, Tongyi Embedding Vision Plus is available through Alibaba Cloud Model Studio and is part of the provider's multimodal embedding offering. It should be viewed as a retrieval and representation model, not as a conversational member of a general-purpose language-model lineup.

How the model represents content

The model returns independent embeddings for submitted content. Its documented embedding size is 1,152 dimensions. In practical terms, each text, image, video, or supported image sequence is represented as a list of numbers. A vector database or search system can then index those values and compare new queries against stored content.

Because the model produces embeddings rather than generated media, its output is not a paragraph, caption, image, audio file, or video. The embedding is an intermediate representation that another application uses for similarity search, ranking, classification, recommendation, or semantic indexing.

Supported input types

Tongyi Embedding Vision Plus accepts Chinese and English text, JPG, PNG, and BMP images, and video supplied through a public URL. Documented video formats include MP4, MPEG, AVI, MOV, MPG, WEBM, FLV, and MKV. The model can also process multi-image sequences containing up to eight images, which can be useful when a single visual item needs to be represented through several frames or views.

The available modalities are therefore broader than those of a text-only embedding model: text, still images, videos, and multi-image inputs can participate in the same retrieval strategy. The model supports image and video input, but it does not provide image or video output.

Input limits and practical requirements

The documented text limit is 1,024 tokens per request. Each image can be up to 3 MB, and a multi-image request can contain no more than eight images. Video input is limited to 10 MB and must be supplied through a public URL.

These restrictions matter when preparing a production indexing pipeline. Long descriptions may need to be divided into shorter records before embedding. Images may need to be resized or compressed to meet the per-image limit, and videos may require preprocessing or hosting at a publicly accessible location before they can be submitted. The supplied research does not specify a context window beyond the 1,024-token text-input limit, and it does not document a generated-output-token limit because the model does not generate text.

What is Tongyi Embedding Vision Plus used for?

The model is most useful when the application needs to find or group semantically related content across media types. Supported use cases include:

  • Text-to-image search: match a natural-language product or scene description to stored images.
  • Image-to-image similarity: find visually or semantically related catalog items, photographs, or reference images.
  • Text-to-video retrieval: locate videos associated with a written query.
  • Video-to-video matching: identify similar clips or related media in a video collection.
  • Multimodal recommendation: use descriptions, product images, and video content as signals for ranking recommendations.
  • Semantic classification: organize content according to its represented meaning rather than relying only on manually assigned keywords.
  • Media indexing: create searchable representations for e-commerce catalogs, galleries, security collections, and autonomous-driving datasets.

A typical implementation would embed the content being indexed, store the resulting vectors, and embed each incoming search query in the same representation space. A vector search system can then return the closest candidates. The model itself supplies the representations; retrieval, filtering, ranking, and application-specific business logic remain outside the model.

Main strengths and trade-offs

The central strength of Tongyi Embedding Vision Plus is its cross-modal scope. A text-only embedding model is appropriate when all of the data is language, but this model is designed for collections where text, images, videos, or image sequences need to be compared in a unified workflow. Its 1,152-dimensional output provides a fixed-size representation that can be indexed consistently across supported input types.

Its other practical advantage is that embedding output is free according to the supplied pricing information. The international Singapore deployment is priced at $0.09 per 1 million input tokens for text, image, and video, while embedding output has a listed price of $0. This makes it positioned toward relatively inexpensive batch indexing and similarity-search pipelines, although the total system cost can still include storage, vector search, networking, media processing, and hosting.

The trade-off is specialization. The model does not answer questions, write code, summarize videos in natural language, call external tools, or generate media. A separate generative model or application component is needed when retrieval results must be explained, transformed, or used to produce a final response.

Pricing, speed, and capability positioning

The documented international Singapore price is $0.09 per 1 million input tokens for text, image, and video. Embedding output is free. The supplied research does not provide a separate monthly subscription, annual commitment, or alternative regional price, so those details should not be assumed.

The research includes editorial comparative scores of 8 for speed and cost, and 1 for reasoning and coding. These are assessments for an embedding model, not provider-published benchmark results. The speed and cost scores reflect the model's focused representation task and listed pricing; they do not establish a guaranteed latency or throughput. Likewise, the low reasoning and coding scores indicate that those are not meaningful purposes for this model, not that the model has been benchmarked as a general reasoning or programming system.

No streaming mode is documented for this exact model. That is consistent with an embedding workflow, where the application generally waits for a vector rather than receiving a progressively generated response. The supplied research also records no support for tool use, function calling, web search, structured outputs, context caching, batch API access, or fine-tuning for this model.

Limitations to consider

Tongyi Embedding Vision Plus is not a conversational or generative model. It returns vectors only, so it cannot directly provide a natural-language explanation of why two items are similar. Applications that need descriptions, question answering, captions, or generated recommendations will need another model after retrieval.

The input constraints may also affect ingestion design. Text is limited to 1,024 tokens per request, images are limited to 3 MB each, multi-image requests are limited to eight images, and videos must be no larger than 10 MB and accessible through a public URL. The model's documented language coverage is Chinese and English; the supplied research does not establish broader language support.

There is also no documented fine-tuning option for this exact model. Organizations with highly specialized visual terminology or domain-specific similarity requirements should evaluate whether the standard representation is sufficient before committing to a large indexing project.

When to choose Tongyi Embedding Vision Plus

Choose this model when the main problem is retrieving or comparing mixed media. It is a strong fit for a product search system that combines descriptions and images, a video archive that needs text and visual retrieval, or a recommendation pipeline that must represent several content types consistently. Its fixed 1,152-dimensional vectors and free embedding output can also be useful for cost-conscious indexing projects, subject to the provider's current pricing and deployment terms.

A dedicated text embedding model such as Alibaba Cloud's text-embedding-v4 may be more appropriate for a text-only workload, especially when image and video support is unnecessary. A conversational or multimodal generative model is more suitable when the application must produce answers, captions, summaries, code, or other user-facing content. In short, Tongyi Embedding Vision Plus should be selected for the quality and convenience of multimodal retrieval, not for reasoning, generation, or tool-based automation.

Bottom line

Tongyi Embedding Vision Plus is a focused Alibaba Cloud embedding model for connecting Chinese and English text with images, videos, and multi-image sequences. Its 1,152-dimensional vectors support cross-modal search, similarity matching, classification, recommendation, and semantic indexing. The model's low listed input price and free embedding output are attractive for retrieval infrastructure, while its lack of generated output, tools, fine-tuning, and conversational capabilities makes its role clear: it is a representation layer for multimodal search, not a general-purpose AI assistant.


Answers to Frequently Asked Questions

What is Tongyi Embedding Vision Plus used for?
Tongyi Embedding Vision Plus is used for cross-modal semantic search and similarity matching across Chinese and English text, images, videos, and multi-image sequences. Common applications include text-to-image search, text-to-video retrieval, image and video matching, recommendation systems, classification, and media indexing.
What input types and limits does Tongyi Embedding Vision Plus support?
The model supports Chinese and English text, JPG, PNG, and BMP images, videos supplied through public URLs, and multi-image sequences of up to eight images. Text is limited to 1,024 tokens per request, each image can be up to 3 MB, and video input can be up to 10 MB. Documented video formats include MP4, MPEG, AVI, MOV, MPG, WEBM, FLV, and MKV.
What does Tongyi Embedding Vision Plus return?
The model returns a 1,152-dimensional numerical embedding for each supported input. These vectors can be stored in a vector database and compared for semantic similarity, retrieval, ranking, classification, recommendation, or indexing. The model does not return text, captions, images, audio, or video.
How much does Tongyi Embedding Vision Plus cost?
According to the supplied pricing information, the international Singapore deployment costs $0.09 per 1 million input tokens for text, image, and video. Embedding output has a listed price of $0. The total cost of a production system may also include storage, vector search, networking, media processing, and hosting.
When should you choose Tongyi Embedding Vision Plus instead of a generative model?
Choose Tongyi Embedding Vision Plus when the primary requirement is retrieving or comparing mixed media such as text, images, and videos. A conversational or multimodal generative model is more appropriate when the application must answer questions, generate captions or summaries, write code, explain results, or create media.


Sources 5
Provider

About Qwen