Seed1.5

Seed1.5-VL

by ByteDance Seed · Retired; Volcano Engine service ended on 2026-03-31

Seed1.5-VL was ByteDance Seed’s multimodal model for image and video understanding, OCR, visual reasoning, grounding, GUI interaction, and gameplay tasks. It combined a 532M vision encoder with a 20B-active-parameter MoE language model and supported tool calling through Volcano Engine. The model’s API was retired on March 31, 2026.

Text Reasoning Coding
Seed1.5-VL was released by ByteDance Seed on May 13, 2025, as a general-purpose model for image and video understanding. It accepted text, images, and video and was designed for tasks ranging from OCR and chart interpretation to visual grounding, long-chain reasoning, GUI interaction, and gameplay analysis. Its Volcano Engine API supported multimodal chat completions, streaming, thinking controls, and function tools. However, the model was retired from Volcano Engine on March 31, 2026, with ByteDance recommending migration to Doubao Seed 2.0 Lite.
Outputs

What Seed1.5-VL can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Streaming Batch API
Model profile

Performance characteristics

8/10 Reasoning
7/10 Coding
6/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Seed1.5
Model type Multimodal
Context window 131K tokens
Release date 2025-05-13
Status Retired; Volcano Engine service ended on 2026-03-31
Deprecation date 2025-12-19
Shutdown date 2026-03-31
Knowledge cutoff notes

ByteDance did not publish a directly verifiable knowledge-cutoff date for Seed1.5-VL in the authoritative model materials reviewed.

Model notes

Seed1.5-VL was the research name for the Volcano Engine model doubao-1-5-thinking-vision-pro-250428. The model used a 532M-parameter vision encoder and a Mixture-of-Experts language model with 20B active parameters. ByteDance reported state-of-the-art results on 38 of 60 public benchmarks. The official retirement notice gave March 31, 2026 as the Volcano Engine end-of-service date and recommended migration to doubao-seed-2-0-lite-260215. The 131,072-token context value comes from the technical report's maximum sequence-length training configuration; an exact retired API max-output limit was not verified.

Model guide

Seed1.5-VL: ByteDance’s Vision-Language Model for Visual Reasoning

Seed1.5-VL was ByteDance Seed’s multimodal vision-language model for understanding images and video, performing visual reasoning, reading documents, locating objects, interpreting diagrams, and supporting GUI-agent and gameplay tasks. It combined a 532-million-parameter vision encoder with a Mixture-of-Experts language model containing 20 billion active parameters. The model was available through Volcano Engine under the identifier doubao-1-5-thinking-vision-pro-250428, but its service ended on March 31, 2026, making it a historical rather than deployable model.

What Seed1.5-VL was designed to do

Seed1.5-VL was ByteDance Seed’s vision-language foundation model for tasks that require an AI system to connect visual information with language-based reasoning. Unlike a text-only language model, it could analyze images and video alongside written instructions, answer questions about visual content, identify relevant regions, interpret documents, and reason through visual problems.

The model was intended for general-purpose multimodal understanding rather than image or video generation. A user could provide a photograph and ask for a description, submit a chart and request an explanation, provide a video and ask what happened at a particular point, or give a screenshot and ask the model to identify an interface element. Its broader target was visual reasoning: solving problems that require several steps instead of simply naming objects in an image.

Seed1.5-VL was released on May 13, 2025. The public research name and the hosted API name were different: Volcano Engine exposed it using the identifier doubao-1-5-thinking-vision-pro-250428.

Architecture and technical profile

According to ByteDance’s technical materials, Seed1.5-VL used a 532-million-parameter SeedViT vision encoder together with a Mixture-of-Experts, or MoE, language model containing 20 billion active parameters. An adapter connected the visual representation produced by the vision encoder to the language model’s multimodal token space.

In an MoE model, different input tokens can be routed through different expert components rather than activating every parameter for every token. The reported active-parameter figure therefore describes the portion used during processing, not necessarily the model’s total stored parameter count. This design allowed Seed1.5-VL to combine a comparatively compact vision component with a much larger reasoning and language-processing system.

The model was pretrained on approximately 3 trillion source tokens. Its visual pipeline supported images at different resolutions through native-resolution transformation. For video, it used dynamic frame-rate and resolution sampling. Timestamp tokens were added before video frames, giving the model explicit temporal information that could help it relate visual events to points in time.

Supported inputs and outputs

Seed1.5-VL accepted three primary input types:

  • Text instructions and questions
  • Still images
  • Video

Its output was text. It did not natively generate images, video, audio, or other media. This distinction matters because the model was multimodal on the input side, but its practical role was to analyze and explain media in language.

CapabilitySupported or reported status
Text inputYes
Image inputYes
Video inputYes
Audio inputNot reported
Text outputYes
Image, video, or audio outputNo
Function and tool useSupported through the Volcano Engine API

The technical report specified a maximum sequence-length training configuration of 131,072 tokens. That value is a context-length reference from the technical materials; it should not be interpreted as a verified current production limit because the hosted service has been retired. An exact maximum output-token limit was not verified.

What Seed1.5-VL could handle

Seed1.5-VL covered a broad set of visual-language tasks:

  • Visual question answering: answering questions about objects, scenes, documents, diagrams, and events shown in an image or video.
  • OCR and document understanding: reading visible text and using it as part of a larger explanation or reasoning task.
  • Charts and diagrams: extracting information from visual representations and describing relationships between their components.
  • Visual grounding: connecting words or descriptions to particular regions, objects, or locations in an image.
  • Counting and localization: finding objects and estimating their positions or quantities.
  • Three-dimensional spatial understanding: reasoning about spatial relationships represented in visual input.
  • Video comprehension: tracking events across frames and relating them to temporal information.
  • Visual puzzles: working through image-based problems that require multiple reasoning steps.
  • GUI interaction and gameplay: interpreting screens and supporting agent-oriented tasks involving interfaces or game environments.

ByteDance reported state-of-the-art results on 38 of 60 public vision-language benchmarks in its release materials, including 14 of 19 video benchmarks and 3 of 7 evaluated GUI-agent tasks. These are provider-reported benchmark claims rather than an independent assessment, and benchmark performance does not guarantee reliable behavior on every real-world image, video, or interface.

Reasoning, coding, and tool support

Reasoning was a central part of Seed1.5-VL’s positioning. The model was built for long-chain multimodal reasoning, meaning it could combine visual evidence with several intermediate steps before producing an answer. Examples include solving a visual puzzle, interpreting a diagram and applying its information to a question, or examining a screenshot to determine which interface element should be used.

Its reasoning ability was primarily visual and language-based rather than a dedicated software-engineering specialization. It could interpret code shown in screenshots or combine visual input with programming-related instructions, but the supplied research does not establish a separate coding benchmark profile or a specialized code-generation mode.

Through the Volcano Engine API, the model supported function and tool calling. This allowed an application to define operations that the model could request during a conversation. Tool support was useful for agent experiments, GUI workflows, and systems that needed to connect visual interpretation with external actions. The model itself did not become an autonomous computer operator simply because tools were available; the surrounding application still had to implement permissions, execution, and safety controls.

Strengths and limitations

Seed1.5-VL’s main strength was the breadth of its visual reasoning target. It was not limited to image captions or basic visual question answering. Its reported capabilities covered images, video, OCR, diagrams, localization, temporal understanding, GUI interaction, and gameplay. The combination of a dedicated vision encoder and an MoE language model was intended to support both detailed visual processing and complex language-based reasoning.

The model also offered a relatively broad context configuration, with the technical report documenting up to 131,072 tokens in its maximum sequence-length training setup. That could be useful for long prompts, extended visual-language conversations, or tasks involving substantial textual context. Because the service is no longer available, however, this figure cannot be used to plan a new production deployment.

ByteDance identified several weaknesses. Fine-grained visual perception could be unreliable, particularly when the task depended on subtle differences between objects or images. Object counting, complex spatial relations, image-difference recognition, combinatorial search, mazes, and sliding puzzles were also listed as difficult areas. Some forms of temporal video reasoning remained challenging, and difficult multi-step tasks could produce unsupported assumptions or incomplete answers.

These limitations are important for practical use. A model may correctly understand the broad content of a photograph while still miscounting similar objects, overlooking a small visual distinction, or making an incorrect assumption about the order of events in a video. Applications that depend on exact localization, safety-critical interpretation, or reliable planning should validate outputs rather than treating the model’s explanation as ground truth.

API availability and lifecycle

Seed1.5-VL was hosted through Volcano Engine’s Ark API. The documented interface supported multimodal chat completions, streaming responses, a thinking-control parameter, and function tools. Its API identity was doubao-1-5-thinking-vision-pro-250428, not simply the research name Seed1.5-VL.

The model is no longer a current deployment option. Volcano Engine’s lifecycle documentation listed March 31, 2026 as the end-of-service date, and ByteDance recommended migration to doubao-seed-2-0-lite-260215. The supplied research also records a deprecation date of December 19, 2025. Since the endpoint has ended, current users should not build a new integration around the old model identifier.

No verified input or output pricing was supplied for Seed1.5-VL, and no current pricing should be inferred from historical API availability. The model’s speed and cost characteristics therefore cannot be used as a present purchasing comparison. For a new project, the relevant evaluation is whether the recommended successor or another currently supported vision-language model offers the required visual accuracy, context capacity, latency, and cost.

When Seed1.5-VL was a suitable choice

When it was available, Seed1.5-VL was a reasonable fit for research and evaluation involving multimodal reasoning rather than media generation. Suitable use cases included:

  • Testing visual question answering and long-form image reasoning
  • Analyzing videos and asking about events or temporal relationships
  • Extracting and explaining text from documents, screenshots, charts, or diagrams
  • Investigating visual grounding, counting, and localization
  • Prototyping GUI-agent or gameplay-oriented systems
  • Comparing multimodal benchmark performance across vision-language models

It was less appropriate for native image or video creation, audio interaction, guaranteed pixel-level accuracy, difficult combinatorial planning, or any current production system requiring a supported endpoint. A faster or less expensive vision model could be preferable for simple classification or routine OCR, while a newer supported reasoning model may be preferable for production workflows that require active maintenance and a current service-level commitment.

Bottom line

Seed1.5-VL was a technically broad vision-language model focused on understanding images and video through language-based reasoning. Its 532-million-parameter vision encoder, 20-billion-active-parameter MoE language model, video processing design, and reported benchmark results made it notable within ByteDance Seed’s 2025 multimodal research portfolio. Its strongest areas were visual reasoning, document and diagram analysis, video understanding, grounding, and agent-oriented visual tasks.

Its practical status is now the decisive consideration: the Volcano Engine service ended on March 31, 2026. Seed1.5-VL remains relevant for understanding ByteDance’s multimodal model development and for historical benchmark comparisons, but new applications should use a currently supported alternative, including the migration target identified by ByteDance, rather than the retired API.


Answers to Frequently Asked Questions

Is Seed1.5-VL still available for new applications?
No. The Volcano Engine service ended on March 31, 2026, so new applications should not be built around the retired identifier. ByteDance recommended migrating to "doubao-seed-2-0-lite-260215" or evaluating another currently supported vision-language model.
What were Seed1.5-VL’s main strengths and limitations?
Its strengths included visual question answering, OCR and document understanding, chart and diagram analysis, video comprehension, visual grounding, localization, and agent-oriented GUI tasks. Reported limitations included fine-grained perception, exact object counting, complex spatial relationships, image-difference recognition, combinatorial search, mazes, sliding puzzles, and some forms of temporal video reasoning.
What inputs and outputs did Seed1.5-VL support?
Seed1.5-VL accepted text, still images, and video as inputs. Its output was text. It did not natively generate images, video, or audio, and audio input was not reported as supported.
What was the Volcano Engine API identifier for Seed1.5-VL?
Volcano Engine exposed Seed1.5-VL through the API identifier "doubao-1-5-thinking-vision-pro-250428". The API supported multimodal chat completions, streaming responses, a thinking-control parameter, and function and tool calling.
What was Seed1.5-VL designed to do?
Seed1.5-VL was ByteDance Seed’s vision-language model for analyzing images and video with language-based reasoning. It could answer visual questions, interpret documents, charts, and diagrams, identify regions or objects, understand video events, and support visual puzzles and GUI-oriented tasks.


Sources 7
Provider

About ByteDance Seed