Seed1.5

Seed1.5 (Doubao-1.5-pro)

by ByteDance Seed · Retired; the primary doubao-1-5-pro-32k-250115 deployment was scheduled to shut down on 2026-09-21 at 14:00 China Standard Time.

Seed1.5 (Doubao-1.5-pro) is ByteDance Seed’s January 2025 multimodal sparse MoE model with text, image, and speech capabilities. Its 32K Volcano Engine deployment supported chat, streaming, tools, and multimodal interaction, but it has been scheduled for retirement in September 2026.

Text Speech Reasoning Coding
Seed1.5 (Doubao-1.5-pro) is ByteDance Seed’s January 2025 general-purpose multimodal foundation model. It uses a sparse Mixture-of-Experts architecture and supports text, visual understanding, and speech interaction, with a 32K-class API deployment through Volcano Engine. The model was designed for knowledge work, coding, reasoning, document understanding, and natural voice interaction. However, its main doubao-1-5-pro-32k-250115 deployment is listed for service shutdown on September 21, 2026.
Outputs

What Seed1.5 (Doubao-1.5-pro) can produce

Text Speech
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Seed1.5
Model type Multimodal
Context window 33K tokens
Release date 2025-01-22
Status Retired; the primary doubao-1-5-pro-32k-250115 deployment was scheduled to shut down on 2026-09-21 at 14:00 China Standard Time.
Deprecation date 2026-07-10
Shutdown date 2026-09-21T14:00:00+08:00
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was found. Web search or provider tools do not change the model's underlying knowledge cutoff.

Model notes

The canonical ByteDance Seed portfolio name is Seed1.5 (Doubao-1.5-pro); the principal Volcano Engine deployment identifier was doubao-1-5-pro-32k-250115. ByteDance described sparse MoE inference, visual understanding, and an end-to-end speech-to-speech framework. A related Doubao-1.5-pro-AS1-Preview reasoning configuration was mentioned in the launch material but is not treated as a separate canonical model. Volcano Engine documentation listed the deployment for retirement with Doubao Seed 2.0 Lite as the migration target. Exact active and total parameter counts, knowledge cutoff, maximum output-token limit, current token prices, fine-tuning, prompt caching, and batch support were not verified from authoritative model-specific documentation.

Model guide

Seed1.5 (Doubao-1.5-pro): ByteDance’s Multimodal MoE Model for Vision and Voice

Seed1.5, officially presented by ByteDance Seed as Doubao-1.5-pro, is a general-purpose sparse Mixture-of-Experts model with text, image, and speech capabilities. Released on January 22, 2025, it combines language understanding, visual comprehension, speech interaction, reasoning, and efficient inference, but its primary Volcano Engine deployment has been scheduled for retirement.

What is Seed1.5 (Doubao-1.5-pro)?

Seed1.5 is a general-purpose multimodal foundation model from ByteDance Seed. On the provider’s model pages and API documentation, it is identified as Doubao-1.5-pro. The two names describe the same canonical model from different perspectives: Seed1.5 is the research and model-family name, while Doubao-1.5-pro is the API-oriented product name.

ByteDance released the model on January 22, 2025. It was positioned as a model for broad language tasks rather than a narrowly specialized image, coding, or speech system. Its supported use cases include question answering, programming assistance, reasoning, document and chart interpretation, visual question answering, and voice-based interaction.

The model is part of ByteDance Seed’s broader foundation-model catalog. It should not be confused with later or related entries such as Seed1.5-VL, Seed1.6, Seed1.8, Seed2.0, or Seed2.1. Those are separate models, even though they belong to the same general ByteDance ecosystem.

Sparse MoE architecture and inference efficiency

Seed1.5 uses a sparse Mixture-of-Experts, or MoE, architecture. In an MoE model, different parts of the network act as specialist modules, and the system activates only a subset of them for each input. This can reduce the computation required for an individual request compared with activating every parameter in a similarly sized dense model.

ByteDance described the model as combining strong general-purpose capability with efficient inference. Its published material discusses separate optimization for prompt processing and token generation, along with techniques such as quantization, speculative decoding, batching, and distributed inference. These techniques are intended to improve time to first token, generation speed, throughput, and serving cost.

The exact active and total parameter counts were not verified in the supplied model-specific documentation. Consequently, the model should be evaluated by its documented behavior and deployment characteristics rather than by an assumed parameter count.

Text, image, and speech capabilities

Seed1.5 is not limited to text. Its documented multimodal capabilities include text input and output, image understanding, document recognition, visual reasoning, and speech interaction.

  • Text: General conversation, knowledge tasks, reasoning, writing, and coding.
  • Images: Image understanding, fine-grained visual interpretation, document analysis, chart interpretation, and visual question answering.
  • Speech: Speech understanding and spoken responses through an end-to-end speech-to-speech interaction framework.

The visual system was described as using native dynamic resolution. This approach is intended to retain useful details across high-resolution images, small images, and images with unusual aspect ratios. That makes the model relevant to tasks such as reading documents, examining charts, and interpreting visual details where a simple low-resolution image representation may lose information.

ByteDance also described an integrated speech-to-speech framework that combines speech and text representations. The intended experience includes understanding spoken input, producing spoken responses, and maintaining expressive or emotional continuity. This is broader than treating speech as a separate transcription or text-to-speech step, although the exact behavior depends on the endpoint and surrounding Volcano Engine service configuration.

Seed1.5 is a multimodal understanding and interaction model, not an image-generation or video-generation model. The supplied research does not verify image, video, or music generation output for this model.

Reasoning and coding performance

ByteDance evaluated Doubao-1.5-pro on knowledge, coding, reasoning, Chinese-language, and multimodal benchmarks. The available research confirms the evaluation areas but does not provide a complete set of benchmark scores, so specific numerical performance claims should not be inferred.

For everyday reasoning, the model was intended to handle multi-step questions, explanations, document interpretation, and general problem solving. It can also support programming assistance, including code generation and technical discussion. These capabilities make it suitable for developers who need a broad assistant that can work with both written instructions and visual material.

ByteDance separately discussed a deep-thinking configuration called Doubao-1.5-pro-AS1-Preview. That configuration was presented as a reasoning-oriented preview and should not be treated as a separate canonical Seed1.5 model record. Users should also avoid assuming that results from the preview configuration automatically describe the standard Doubao-1.5-pro deployment.

Context window, API access, and tools

The primary documented deployment was doubao-1-5-pro-32k-250115 on Volcano Engine’s ModelArk platform. Its documented context length is 32,768 tokens. This is enough for many conversations, source files, reports, and moderately long documents, although the usable amount may also depend on system instructions, tool messages, multimodal content, and the provider’s request rules.

The API supported chat completions, streaming responses, function or tool calls, and multimodal requests through the provider’s platform. Tool-enabled bot integrations could expose functions such as search or other platform actions. Those tools belong to the surrounding provider service layer; they do not mean that the model has an automatically updated internal knowledge base.

A maximum output-token limit was not verified in the supplied documentation. Fine-tuning, prompt caching, batch processing, and a dedicated JSON mode were also not confirmed for this model. These should remain open questions when planning a production integration rather than being assumed from features available in newer or related ByteDance models.

Pricing and deployment lifecycle

No authoritative model-specific input or output token prices were verified in the supplied research. Pricing may depend on the Volcano Engine account, region, endpoint, billing arrangement, or historical availability. The model should therefore not be compared on exact per-token cost without checking the applicable ModelArk documentation and account terms.

Lifecycle status is an important practical limitation. Volcano Engine documentation identified doubao-1-5-pro-32k-250115 for retirement and named Doubao Seed 2.0 Lite as the migration target. The supplied lifecycle information gives a scheduled service shutdown of September 21, 2026, at 14:00 China Standard Time. The model record is consequently best treated as a historical or migration-bound deployment rather than a dependable choice for a new long-lived production system.

The supplied database also records a deprecation date of July 10, 2026. Because the available lifecycle sources distinguish deprecation from final shutdown, users should consult the latest Volcano Engine notice for the precise availability state and any changes to the schedule.

Main strengths and limitations

Strengths

  • Broad modality coverage: The model combines text, image understanding, and speech interaction instead of requiring separate models for every input type.
  • Useful document and visual analysis: Dynamic-resolution visual processing was designed for detailed image, document, and chart interpretation.
  • Integrated voice interaction: The speech-to-speech design supports spoken conversations rather than only transcription or text responses.
  • General-purpose coverage: Knowledge work, coding, reasoning, and multimodal tasks are handled within one model family.
  • Serving efficiency: Sparse MoE execution and provider-described inference optimizations target a balance between quality, speed, and operating cost.

Limitations

  • Retirement risk: The main documented deployment has a scheduled shutdown, making it unsuitable for new systems that require long-term model continuity.
  • Incomplete public specifications: Exact parameter counts, maximum output tokens, knowledge cutoff, current pricing, fine-tuning, caching, and batch support were not verified.
  • Endpoint dependence: Multimodal and speech features may vary according to the specific Volcano Engine endpoint and service layer.
  • No verified generation role: The research supports image and document understanding, not image, video, or music generation.
  • Regional and platform constraints: Access and feature availability depend on ByteDance’s developer platforms, account arrangements, and applicable regional policies.

When to choose this model

Seed1.5 is most appropriate when an organization needs a historically significant ByteDance multimodal model for evaluation, migration analysis, or an existing Volcano Engine integration. It is a reasonable fit for applications that combine ordinary chat with image or document understanding, coding assistance, and voice interaction.

Its sparse MoE design may also appeal to teams balancing response quality against inference efficiency. In the supplied editorial evaluation, it received relatively strong assessments for speed and cost compared with its broader reasoning and multimodal capabilities. Those are editorial scores, not provider-published benchmark facts, and should be validated against the specific endpoint and workload.

A different option is more appropriate for a new production deployment if the project needs a currently supported model, guaranteed lifecycle stability, transparent token pricing, or modern structured-output features. A dedicated vision model may be preferable for specialized image analysis, while a current reasoning model may be a better choice for demanding long-horizon problem solving. For ByteDance users, the documented migration target, Doubao Seed 2.0 Lite, should be evaluated where continued platform support is more important than preserving the exact Seed1.5 behavior.

Bottom line

Seed1.5 (Doubao-1.5-pro) is a general-purpose ByteDance model that brought sparse-MoE efficiency, visual understanding, and speech interaction together in a 32K-class deployment. Its strongest distinction is the combination of language, image, document, and voice capabilities in one model rather than specialization in a single modality.

Its principal practical drawback is lifecycle status. With the primary deployment scheduled for retirement and several important specifications left undocumented, Seed1.5 is better understood as an existing or historical model to evaluate carefully than as the default foundation for a new long-term application.


Answers to Frequently Asked Questions

Is Seed1.5 still suitable for a new production deployment?
Seed1.5 should be approached cautiously for new long-term systems because Volcano Engine has identified doubao-1-5-pro-32k-250115 for retirement. The documented service shutdown is September 21, 2026, at 14:00 China Standard Time, with Doubao Seed 2.0 Lite named as the migration target.
What is the context window and primary API deployment for Seed1.5?
The primary documented deployment is doubao-1-5-pro-32k-250115 on Volcano Engine’s ModelArk platform. It provides a documented context length of 32,768 tokens and supports chat completions, streaming, function or tool calls, and multimodal requests.
What is Seed1.5 (Doubao-1.5-pro)?
Seed1.5 is a general-purpose multimodal foundation model from ByteDance Seed, identified in API documentation as Doubao-1.5-pro. It supports text, image and document understanding, reasoning, coding assistance, and speech-based interaction.
What modalities and capabilities does Seed1.5 support?
Seed1.5 supports text input and output, image understanding, document and chart interpretation, visual question answering, speech understanding, and spoken responses through an end-to-end speech-to-speech framework. It is not verified as an image-, video-, or music-generation model.


Sources 6
Provider

About ByteDance Seed