What is Seed-1.6-Embedding?
Seed-1.6-Embedding is a multimodal embedding model developed by ByteDance’s Seed team and launched on June 28, 2025. It is based on Seed1.6-Flash, but its purpose is different from that of a chat or text-generation model: it converts content into numerical vectors that represent semantic meaning.
These vectors, commonly called embeddings, can be stored in a vector database and compared with other vectors. Items with similar meaning tend to be located near one another in the resulting representation space. For example, a search query such as “a red sports car on a mountain road” can be compared with image vectors to retrieve visually and semantically related images. An image can likewise be used to find related text, images, or video content.
The model is exposed through Volcano Engine with the API model ID doubao-embedding-vision-250615. Its documented role is developer and enterprise-oriented vectorization, not direct conversation or content generation.
Supported inputs and output
Seed-1.6-Embedding accepts text, images, video, and combinations of these modalities. Its main distinction is that these inputs can be mapped into a shared semantic space, enabling cross-modal retrieval rather than requiring a separate search model for every content type.
- Text input: Suitable for search queries, captions, documents, labels, and other textual content.
- Image input: Can be indexed or compared with text and other visual content.
- Video input: Supports video retrieval and indexing through sampled video frames.
- Mixed-modal input: Text, images, and video can be used in retrieval workflows where the application needs to combine multiple signals.
The model returns dense floating-point embedding vectors rather than natural-language text, images, audio, or video. The default vector size is 2,048 dimensions, while a 1,024-dimensional option is also available. The smaller representation can reduce storage and retrieval costs when the application does not require the larger vector size.
The API can return embeddings in floating-point form or as Base64-encoded data. Seed-1.6-Embedding also supports sparse embeddings for text-only input according to the documented multimodal vectorization interface. The 250615 version does not support the newer multi-embedding output format.
How cross-modal retrieval works
In a conventional text search system, a text query is matched primarily against text. A multimodal embedding system instead gives different content types compatible representations. A text description and an image showing the described scene can therefore be compared using the same general retrieval mechanism.
Practical examples include:
- Searching a product-image library with a natural-language description.
- Finding documents, captions, or support records related to an uploaded image.
- Locating visually or semantically similar moments in an indexed video collection.
- Combining text metadata and visual content when ranking items in a recommendation system.
- Retrieving images or video clips as context for a separate generative model.
Video is processed through sampled frames. The supplied documentation states that the model does not currently understand audio contained in video files. Applications that need speech, sound effects, or other audio-based video search should therefore extract and process audio separately with an appropriate model.
Technical limits and API characteristics
Volcano Engine documentation lists a 128K context window for the 250615 model version. This is the documented model-level context figure, but it should not be confused with the per-input limits of the multimodal vectorization interface.
Individual text inputs are limited to 8K tokens in the documented interface. Video files are limited to 50 MB. The 250615 version also lacks custom video frame-sampling controls, so applications requiring precise control over which frames are selected may need to preprocess video themselves or use a different supported version or workflow.
The principal output limit is the embedding dimension rather than a generated-token budget: vectors contain either 2,048 dimensions by default or 1,024 dimensions when the smaller option is selected. There is no conversational maximum-output-token setting because the model does not generate prose responses.
| Specification | Seed-1.6-Embedding |
|---|---|
| Provider | ByteDance Seed |
| Volcano Engine model ID | doubao-embedding-vision-250615 |
| Release date | June 28, 2025 |
| Input modalities | Text, images, video, and mixed text-image-video inputs |
| Output | Dense embeddings; 2,048 dimensions by default or 1,024 optionally |
| Context window | 128K, according to Volcano Engine documentation |
| Individual text limit | 8K tokens in the documented multimodal vectorization interface |
| Video file limit | 50 MB |
| Video audio understanding | Not supported |
What is it best used for?
Seed-1.6-Embedding is best suited to systems where content must be found, grouped, ranked, or compared rather than generated. Its shared representation is particularly useful when the query and the stored content use different modalities.
Semantic search and indexing
Organizations can embed documents, images, product records, or video segments and store the vectors in a vector database. Search queries can then retrieve content by meaning instead of relying only on exact keywords. For video, indexing may involve creating representations from sampled frames and associating them with timestamps or metadata in the application.
Multimodal retrieval-augmented generation
A retrieval-augmented generation system can use Seed-1.6-Embedding to find relevant images, text, or video before passing that context to a separate generative model. This is useful for visual knowledge bases, product-support systems, media archives, and applications that need evidence from multiple content types.
Classification, clustering and recommendation
Embedding vectors can also support grouping and similarity-based ranking. Similar products, documents, images, or clips can be clustered together, while recommendations can be based on semantic or visual proximity. These are application-level uses: the model supplies representations, while the surrounding database or machine-learning system determines how items are ranked and organized.
Strengths and trade-offs
The main strength of Seed-1.6-Embedding is its cross-modal design. A single model family can support text-image and video-related retrieval instead of limiting the application to text-only semantic search. The choice between 2,048 and 1,024 dimensions also gives developers a storage and retrieval trade-off.
The model is not a universal AI endpoint. It produces vectors, not answers. It has no documented tool or function-calling capability, streaming generation, structured-output mode, or native speech, image, or video generation. It also is not intended for reasoning conversations or code generation. The supplied model evaluation fields rate reasoning and coding usefulness at 1, but these are editorial database scores, not provider-published benchmark results and should not be interpreted as capabilities of the embedding model.
Speed and cost are also different from those of a chat model. Embeddings are normally generated as an indexing or retrieval service, and the smaller 1,024-dimensional option can reduce the amount of data stored and processed. However, a model-specific public price for Seed-1.6-Embedding was not verified in the supplied official documentation. Volcano Engine usage is described as usage-based and may depend on the deployed service and input consumption, so current pricing should be checked in the relevant Volcano Engine account or product documentation.
When to choose Seed-1.6-Embedding
Choose Seed-1.6-Embedding when the central problem is retrieving or comparing text, images, video, or combinations of these. It is a strong fit for a media archive, visual product search, multimodal knowledge base, recommendation pipeline, or retrieval system that must accept a text query and return non-text content.
The 2,048-dimensional output is the appropriate documented default when the application prioritizes representation capacity and can accommodate the associated vector-storage cost. The 1,024-dimensional option is worth considering when database size, memory, or retrieval efficiency is more important. The choice should be validated against the application’s own search-quality requirements rather than assumed from dimension count alone.
Another option may be more appropriate when the application needs direct natural-language answers, code generation, tool use, speech understanding, video-audio analysis, or generated images and video. A separate generative model can be combined with Seed-1.6-Embedding in a retrieval-augmented system, but Seed-1.6-Embedding itself should be treated as the retrieval and representation component.
Availability and important limitations
Seed-1.6-Embedding is currently identified as available through Volcano Engine. Access, regional availability, account requirements, quotas, and final service pricing may depend on the selected Volcano Engine deployment and are not fully specified in the supplied research.
Before adopting it, developers should test representative content from their own domain. In particular, they should evaluate the effect of sampled-frame video processing, the 8K-token text-input limit, the 50 MB video limit, and the absence of video-audio understanding. They should also decide whether 1,024-dimensional vectors provide sufficient retrieval quality or whether the default 2,048-dimensional output is justified by the application’s storage and search budget.

