Seed1.6

Seed-1.6-Embedding

by ByteDance Seed · Current and available through Volcano Engine

Seed-1.6-Embedding is ByteDance’s multimodal vectorization model for text, image, video, and mixed-content retrieval. The review covers its Volcano Engine model ID, shared embedding space, vector dimensions, context and input limits, video-processing restrictions, use cases, pricing availability, and trade-offs against generative models.

Embeddings Reasoning Coding
Seed-1.6-Embedding is a vectorization model rather than a conversational AI assistant. Built on Seed1.6-Flash, it represents text, images, videos, or mixed inputs as dense embeddings in a shared semantic space. This allows an application to search for images with text, find text related to an image, index video content, or compare multimodal items using one retrieval system. The model is available through Volcano Engine under the model ID doubao-embedding-vision-250615.
Outputs

What Seed-1.6-Embedding can produce

Embeddings
Inputs

What it can understand

Text Images Video Multimodal input
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Seed1.6
Model type Other
Context window 128K tokens
Release date 2025-06-28
Status Current and available through Volcano Engine
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was identified. Embedding models are generally evaluated through their trained representation capabilities rather than a published conversational knowledge cutoff.

Model notes

Seed-1.6-Embedding was launched on June 28, 2025 and is exposed through Volcano Engine as doubao-embedding-vision-250615. It supports text, image, video, and mixed-modal input in a shared embedding space. Dense output dimensions are 2048 by default or 1024 optionally. The 250615 version supports sparse embeddings for text-only input, but not multi-embedding output or custom video frame sampling. Official documentation states that individual text input is limited to 8K tokens in the multimodal vectorization interface, the overall model context window is 128K, and individual video files are limited to 50 MB. Video audio is not understood. A model-specific public price was not verified.

Model guide

Seed-1.6-Embedding: ByteDance’s Cross-Modal Vector Model for Text, Image and Video Search

Seed-1.6-Embedding is a ByteDance multimodal embedding model for cross-modal retrieval across text, images, and video. Available through Volcano Engine as doubao-embedding-vision-250615, it converts supported inputs into shared vector representations for semantic search, indexing, classification, clustering, recommendation, and multimodal retrieval-augmented generation.

What is Seed-1.6-Embedding?

Seed-1.6-Embedding is a multimodal embedding model developed by ByteDance’s Seed team and launched on June 28, 2025. It is based on Seed1.6-Flash, but its purpose is different from that of a chat or text-generation model: it converts content into numerical vectors that represent semantic meaning.

These vectors, commonly called embeddings, can be stored in a vector database and compared with other vectors. Items with similar meaning tend to be located near one another in the resulting representation space. For example, a search query such as “a red sports car on a mountain road” can be compared with image vectors to retrieve visually and semantically related images. An image can likewise be used to find related text, images, or video content.

The model is exposed through Volcano Engine with the API model ID doubao-embedding-vision-250615. Its documented role is developer and enterprise-oriented vectorization, not direct conversation or content generation.

Supported inputs and output

Seed-1.6-Embedding accepts text, images, video, and combinations of these modalities. Its main distinction is that these inputs can be mapped into a shared semantic space, enabling cross-modal retrieval rather than requiring a separate search model for every content type.

  • Text input: Suitable for search queries, captions, documents, labels, and other textual content.
  • Image input: Can be indexed or compared with text and other visual content.
  • Video input: Supports video retrieval and indexing through sampled video frames.
  • Mixed-modal input: Text, images, and video can be used in retrieval workflows where the application needs to combine multiple signals.

The model returns dense floating-point embedding vectors rather than natural-language text, images, audio, or video. The default vector size is 2,048 dimensions, while a 1,024-dimensional option is also available. The smaller representation can reduce storage and retrieval costs when the application does not require the larger vector size.

The API can return embeddings in floating-point form or as Base64-encoded data. Seed-1.6-Embedding also supports sparse embeddings for text-only input according to the documented multimodal vectorization interface. The 250615 version does not support the newer multi-embedding output format.

How cross-modal retrieval works

In a conventional text search system, a text query is matched primarily against text. A multimodal embedding system instead gives different content types compatible representations. A text description and an image showing the described scene can therefore be compared using the same general retrieval mechanism.

Practical examples include:

  • Searching a product-image library with a natural-language description.
  • Finding documents, captions, or support records related to an uploaded image.
  • Locating visually or semantically similar moments in an indexed video collection.
  • Combining text metadata and visual content when ranking items in a recommendation system.
  • Retrieving images or video clips as context for a separate generative model.

Video is processed through sampled frames. The supplied documentation states that the model does not currently understand audio contained in video files. Applications that need speech, sound effects, or other audio-based video search should therefore extract and process audio separately with an appropriate model.

Technical limits and API characteristics

Volcano Engine documentation lists a 128K context window for the 250615 model version. This is the documented model-level context figure, but it should not be confused with the per-input limits of the multimodal vectorization interface.

Individual text inputs are limited to 8K tokens in the documented interface. Video files are limited to 50 MB. The 250615 version also lacks custom video frame-sampling controls, so applications requiring precise control over which frames are selected may need to preprocess video themselves or use a different supported version or workflow.

The principal output limit is the embedding dimension rather than a generated-token budget: vectors contain either 2,048 dimensions by default or 1,024 dimensions when the smaller option is selected. There is no conversational maximum-output-token setting because the model does not generate prose responses.

SpecificationSeed-1.6-Embedding
ProviderByteDance Seed
Volcano Engine model IDdoubao-embedding-vision-250615
Release dateJune 28, 2025
Input modalitiesText, images, video, and mixed text-image-video inputs
OutputDense embeddings; 2,048 dimensions by default or 1,024 optionally
Context window128K, according to Volcano Engine documentation
Individual text limit8K tokens in the documented multimodal vectorization interface
Video file limit50 MB
Video audio understandingNot supported

What is it best used for?

Seed-1.6-Embedding is best suited to systems where content must be found, grouped, ranked, or compared rather than generated. Its shared representation is particularly useful when the query and the stored content use different modalities.

Semantic search and indexing

Organizations can embed documents, images, product records, or video segments and store the vectors in a vector database. Search queries can then retrieve content by meaning instead of relying only on exact keywords. For video, indexing may involve creating representations from sampled frames and associating them with timestamps or metadata in the application.

Multimodal retrieval-augmented generation

A retrieval-augmented generation system can use Seed-1.6-Embedding to find relevant images, text, or video before passing that context to a separate generative model. This is useful for visual knowledge bases, product-support systems, media archives, and applications that need evidence from multiple content types.

Classification, clustering and recommendation

Embedding vectors can also support grouping and similarity-based ranking. Similar products, documents, images, or clips can be clustered together, while recommendations can be based on semantic or visual proximity. These are application-level uses: the model supplies representations, while the surrounding database or machine-learning system determines how items are ranked and organized.

Strengths and trade-offs

The main strength of Seed-1.6-Embedding is its cross-modal design. A single model family can support text-image and video-related retrieval instead of limiting the application to text-only semantic search. The choice between 2,048 and 1,024 dimensions also gives developers a storage and retrieval trade-off.

The model is not a universal AI endpoint. It produces vectors, not answers. It has no documented tool or function-calling capability, streaming generation, structured-output mode, or native speech, image, or video generation. It also is not intended for reasoning conversations or code generation. The supplied model evaluation fields rate reasoning and coding usefulness at 1, but these are editorial database scores, not provider-published benchmark results and should not be interpreted as capabilities of the embedding model.

Speed and cost are also different from those of a chat model. Embeddings are normally generated as an indexing or retrieval service, and the smaller 1,024-dimensional option can reduce the amount of data stored and processed. However, a model-specific public price for Seed-1.6-Embedding was not verified in the supplied official documentation. Volcano Engine usage is described as usage-based and may depend on the deployed service and input consumption, so current pricing should be checked in the relevant Volcano Engine account or product documentation.

When to choose Seed-1.6-Embedding

Choose Seed-1.6-Embedding when the central problem is retrieving or comparing text, images, video, or combinations of these. It is a strong fit for a media archive, visual product search, multimodal knowledge base, recommendation pipeline, or retrieval system that must accept a text query and return non-text content.

The 2,048-dimensional output is the appropriate documented default when the application prioritizes representation capacity and can accommodate the associated vector-storage cost. The 1,024-dimensional option is worth considering when database size, memory, or retrieval efficiency is more important. The choice should be validated against the application’s own search-quality requirements rather than assumed from dimension count alone.

Another option may be more appropriate when the application needs direct natural-language answers, code generation, tool use, speech understanding, video-audio analysis, or generated images and video. A separate generative model can be combined with Seed-1.6-Embedding in a retrieval-augmented system, but Seed-1.6-Embedding itself should be treated as the retrieval and representation component.

Availability and important limitations

Seed-1.6-Embedding is currently identified as available through Volcano Engine. Access, regional availability, account requirements, quotas, and final service pricing may depend on the selected Volcano Engine deployment and are not fully specified in the supplied research.

Before adopting it, developers should test representative content from their own domain. In particular, they should evaluate the effect of sampled-frame video processing, the 8K-token text-input limit, the 50 MB video limit, and the absence of video-audio understanding. They should also decide whether 1,024-dimensional vectors provide sufficient retrieval quality or whether the default 2,048-dimensional output is justified by the application’s storage and search budget.


Answers to Frequently Asked Questions

What are the main limits of Seed-1.6-Embedding?
The documented limits include an 8K-token limit for individual text inputs, a 50 MB video-file limit, and a 128K model-level context window. The 250615 version does not support video-audio understanding, custom video frame-sampling controls, or the newer multi-embedding output format.
What embedding dimensions does Seed-1.6-Embedding provide?
Seed-1.6-Embedding returns dense floating-point vectors with 2,048 dimensions by default. A 1,024-dimensional option is also available to reduce storage and retrieval costs.
What is the Volcano Engine model ID for Seed-1.6-Embedding?
The Volcano Engine model ID is doubao-embedding-vision-250615.
Which input types does Seed-1.6-Embedding support?
The model supports text, images, video, and combinations of text, images, and video. Video retrieval is based on sampled frames, and the model does not currently understand audio contained in video files.
What is Seed-1.6-Embedding?
Seed-1.6-Embedding is a multimodal embedding model developed by ByteDance’s Seed team. It converts text, images, video, and mixed-modal inputs into numerical vectors for semantic search, indexing, clustering, recommendation, and retrieval-augmented generation.


Sources 5
Provider

About ByteDance Seed