Embedding-Vision

tongyi-embedding-vision-flash

by Qwen · Current and available through Alibaba Cloud Model Studio International deployment

Alibaba Cloud's Tongyi Embedding Vision Flash converts text, images, videos, and image sequences into fixed 768-dimensional vectors. It is designed for cost-sensitive cross-modal retrieval, media catalog indexing, and vector search, with documented Singapore pricing of $0.03 per million input tokens for image or video and $0.09 per million for text.

Embeddings Reasoning Coding
Tongyi Embedding Vision Flash is a vision-centric multimodal embedding model available through Alibaba Cloud Model Studio. It accepts text, images, videos, and multi-image sequences, then returns dense 768-dimensional embeddings that applications can store in a vector database for search, matching, classification, and recommendation workflows. Its fixed output size and lower published input prices make it a practical option when retrieval cost and storage efficiency matter more than using a larger embedding model.
Outputs

What tongyi-embedding-vision-flash can produce

Embeddings
Inputs

What it can understand

Text Images Video Multimodal input
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Embedding-Vision
Model type Multimodal
Context window 1K tokens
Maximum output tokens
Release date 2025-10-31
Status Current and available through Alibaba Cloud Model Studio International deployment
Knowledge cutoff notes

Alibaba Cloud's public documentation specifies the model's input modalities, vector dimensions, limits, and pricing but does not publish a knowledge cutoff for this embedding model.

Model notes

Alibaba Cloud describes this as the lightweight version of its vision-centric Embedding-Vision model family. It returns fixed 768-dimensional dense independent embeddings and does not support the dimension parameter. Supported inputs include text, images, videos, and multi_images sequences. The model supports independent vectors only; it does not support fused embeddings. For the International Singapore deployment, image and video inputs cost $0.03 per 1 million input tokens and text costs $0.09 per 1 million input tokens. The model was listed in Alibaba Cloud's model lifecycle catalog with a 2025-10-31 release date.

Cost

Model pricing

Input Image/video: $0.03 per 1 million input tokens; text: $0.09 per 1 million input tokens
Output Free; embedding output is not charged
Model guide

Tongyi Embedding Vision Flash: Cost-Efficient Multimodal Search Embeddings

Tongyi Embedding Vision Flash is Alibaba Cloud's lightweight multimodal embedding model for converting text, images, videos, and image sequences into fixed 768-dimensional vectors. It is designed for cost-sensitive cross-modal retrieval, catalog indexing, and similarity search rather than text generation or conversational use.

What is Tongyi Embedding Vision Flash?

Tongyi Embedding Vision Flash is a multimodal embedding model from Alibaba Cloud. An embedding is a numerical representation of content: the model converts an input such as a sentence, image, video, or image sequence into a vector that software can compare with other vectors. Similar content should produce vectors that are closer together, allowing applications to build semantic and cross-modal search systems.

The model is part of Alibaba Cloud's vision-centric Embedding-Vision family and is positioned as the lightweight Flash variant. Its output is a fixed 768-dimensional dense vector. The model does not generate prose, images, audio, or video; its job is to represent content for downstream retrieval and analysis.

Alibaba Cloud lists the model as available through Model Studio's International deployment in Singapore. The supplied lifecycle information gives a release date of October 31, 2025, and identifies the model as current and available through that deployment.

Supported inputs and outputs

Tongyi Embedding Vision Flash supports the following input types:

  • Text, with a maximum length of 1,024 tokens.
  • Images, supplied through public URLs or supported Base64 data URIs.
  • Videos supplied through publicly accessible URLs.
  • Multi-image sequences through the multi_images input type.

Documented image formats include JPG, PNG, and BMP. Documented video formats include MP4, MPEG, AVI, MOV, MPG, WEBM, FLV, and MKV. These inputs can be used independently or in supported combinations, but the model returns independent vectors rather than automatically merging all modalities into one fused representation.

The output is a dense 768-dimensional embedding. The dimension is fixed for this model, and the documented interface does not support changing it with a dimension parameter. Embedding output is not charged according to the supplied pricing information.

How the model works in practice

A typical workflow sends a collection of product images, media files, or text descriptions to the model and stores the returned vectors in a vector index. A search query is then embedded in the same representation space. For example, a text query can be compared with image vectors to support text-to-image search, while a reference image can be compared with catalog images for visual similarity.

Supported retrieval patterns include text-to-image, image-to-image, text-to-video, and video-to-video search. The model can also support image classification and multimedia catalog indexing when an application uses embedding similarity or a downstream classifier.

The distinction between independent and fused vectors is important. Independent vectors represent each submitted modality separately. They do not constitute a single vector that combines the meaning of a text description and an image into one fused representation. Applications that specifically require fused multimodal embeddings should evaluate a model designed for that purpose, such as Qwen3-VL-Embedding, rather than assuming that Tongyi Embedding Vision Flash performs fusion.

Technical specifications and limitations

SpecificationDocumented value
Model typeMultimodal embedding model
Embedding dimension768, fixed
Maximum text length1,024 tokens
Input modalitiesText, images, videos, and multi-image sequences
OutputDense independent vector embeddings
Configurable dimensionsNot supported
Text generationNot supported
Image or video generationNot supported

The 1,024-token limit applies to text input. The supplied research does not specify a maximum video duration, maximum image size, maximum number of images in a sequence, or a maximum output-token setting. Those values should not be assumed from the text limit.

The model is not a reranker. An embedding model retrieves or compares candidates by representing them as vectors; a reranker is a separate component that can reorder retrieved results using a more detailed relevance assessment. Tongyi Embedding Vision Flash is also not a conversational model and should not be selected for answer generation or general-purpose assistant tasks.

Pricing

For the International Singapore deployment, Alibaba Cloud lists image and video input at $0.03 per 1 million input tokens and text input at $0.09 per 1 million input tokens. The supplied pricing information states that embedding output is free.

These prices are input-modality-specific, so a workload's cost depends on the mixture of text, images, and videos it processes. The published rates are particularly relevant for large media-indexing jobs, although an application should still account for storage, vector database, network, and any separate retrieval or reranking services.

Pricing and availability can vary by deployment and may change. The figures above describe the documented International Singapore deployment rather than every Alibaba Cloud region.

Strengths and trade-offs

The clearest practical strength is the combination of multimodal input support and a relatively compact fixed vector size. A 768-dimensional vector requires less storage and can reduce the computational burden of similarity search compared with a higher-dimensional representation, especially when a system indexes a large image or video library. This is an engineering trade-off rather than a provider-published benchmark result.

The model also covers several media types within one embedding family. A team building a multimedia catalog can use it for text, images, videos, and image sequences instead of maintaining unrelated representations for every input type. Its published image and video rate is lower than its text rate, which may be useful for media-heavy indexing workloads.

There are corresponding limitations. The output dimension cannot be configured, and independent vectors may not meet the needs of an application that requires a single fused representation across modalities. The model also provides no generative output, reasoning workflow, tool use, or reranking capability. Those functions would need to be supplied by other models or application components.

Reasoning, coding, and tool support

Tongyi Embedding Vision Flash is not a reasoning or coding model in the usual language-model sense. It transforms inputs into embeddings and does not produce explanations, program code, or conversational responses. The supplied model data assigns low reasoning and coding scores, but those are editorial evaluation fields, not claims published by Alibaba Cloud and not benchmark measurements.

The documented model profile does not list function calling, tool use, web search, streaming, structured output, or a batch API. These capabilities should therefore be treated as unsupported or unverified for this model rather than inferred from the broader Model Studio platform. Fine-tuning and caching are also not specified in the supplied research.

When to choose Tongyi Embedding Vision Flash

Choose this model when the main task is multimodal retrieval and the system benefits from a compact, fixed-size vector. Suitable examples include:

  • Searching an image catalog with natural-language descriptions.
  • Finding visually similar products or media assets.
  • Searching video archives by text or comparing related videos.
  • Indexing e-commerce, security, autonomous-driving, or other multimedia datasets.
  • Building a cost-sensitive vector-search system where 768 dimensions are sufficient.

It is especially appropriate when the application needs independent representations for different media types and does not require the embedding model itself to generate answers. Its fixed output can make schema design straightforward and can help control vector storage at large scale.

When another option may be more appropriate

Use a generative language or vision-language model when the requirement is to explain an image, answer questions, write code, summarize a video, or conduct a conversation. Tongyi Embedding Vision Flash cannot perform those tasks.

Use a reranking service when first-stage vector retrieval is not enough and the application needs a separate relevance-ranking step. Use a fused multimodal embedding model when text and visual information must be represented together in one vector. The supplied research specifically identifies Qwen3-VL-Embedding as an example of a more appropriate direction for fused multimodal retrieval.

Finally, a higher-dimensional sibling may be preferable when the application has evidence that additional representation capacity improves its retrieval quality and the associated storage and search costs are acceptable. Alibaba Cloud's Tongyi Embedding Vision Plus is described as returning 1,152-dimensional vectors, compared with Flash's fixed 768 dimensions. The supplied research does not provide benchmark results showing how much quality differs, so the choice should be validated on the application's own data.

Bottom line

Tongyi Embedding Vision Flash is a focused embedding model for cost-conscious multimodal search. Its verified profile is straightforward: text, image, video, and multi-image inputs; a 1,024-token text limit; fixed 768-dimensional dense embeddings; and published Singapore pricing of $0.03 per million input tokens for image or video and $0.09 per million for text. It is a good fit for cross-modal retrieval and large media indexes, but not for generation, conversation, reranking, tool use, or fused multimodal representations.


Answers to Frequently Asked Questions

What are the main limitations of Tongyi Embedding Vision Flash?
The model has a fixed 768-dimensional output and does not support configurable dimensions. It does not generate text, images, or videos; perform conversation, reasoning, or coding; rerank search results; or provide documented tool use, streaming, structured output, or batch API support.
How much does Tongyi Embedding Vision Flash cost?
For the International Singapore deployment, Alibaba Cloud lists image and video input at $0.03 per 1 million input tokens and text input at $0.09 per 1 million input tokens. Embedding output is listed as free. Pricing may vary by deployment and can change.
Does Tongyi Embedding Vision Flash create a fused multimodal embedding?
No. Tongyi Embedding Vision Flash returns independent vectors for submitted modalities rather than one vector that fuses text and visual content. Applications requiring fused multimodal representations should evaluate a model designed for that purpose, such as Qwen3-VL-Embedding.
What input types and output dimensions does Tongyi Embedding Vision Flash support?
The model accepts text, images, videos, and multi-image sequences. Text input supports up to 1,024 tokens, and documented media formats include JPG, PNG, BMP, MP4, MPEG, AVI, MOV, MPG, WEBM, FLV, and MKV. It returns fixed, dense 768-dimensional embeddings.
What is Tongyi Embedding Vision Flash used for?
Tongyi Embedding Vision Flash is used to create embeddings for multimodal search and similarity systems. It supports text-to-image, image-to-image, text-to-video, and video-to-video retrieval, as well as multimedia catalog indexing and image classification workflows.


Sources 5
Provider

About Qwen