What Seed1.5-VL was designed to do
Seed1.5-VL was ByteDance Seed’s vision-language foundation model for tasks that require an AI system to connect visual information with language-based reasoning. Unlike a text-only language model, it could analyze images and video alongside written instructions, answer questions about visual content, identify relevant regions, interpret documents, and reason through visual problems.
The model was intended for general-purpose multimodal understanding rather than image or video generation. A user could provide a photograph and ask for a description, submit a chart and request an explanation, provide a video and ask what happened at a particular point, or give a screenshot and ask the model to identify an interface element. Its broader target was visual reasoning: solving problems that require several steps instead of simply naming objects in an image.
Seed1.5-VL was released on May 13, 2025. The public research name and the hosted API name were different: Volcano Engine exposed it using the identifier doubao-1-5-thinking-vision-pro-250428.
Architecture and technical profile
According to ByteDance’s technical materials, Seed1.5-VL used a 532-million-parameter SeedViT vision encoder together with a Mixture-of-Experts, or MoE, language model containing 20 billion active parameters. An adapter connected the visual representation produced by the vision encoder to the language model’s multimodal token space.
In an MoE model, different input tokens can be routed through different expert components rather than activating every parameter for every token. The reported active-parameter figure therefore describes the portion used during processing, not necessarily the model’s total stored parameter count. This design allowed Seed1.5-VL to combine a comparatively compact vision component with a much larger reasoning and language-processing system.
The model was pretrained on approximately 3 trillion source tokens. Its visual pipeline supported images at different resolutions through native-resolution transformation. For video, it used dynamic frame-rate and resolution sampling. Timestamp tokens were added before video frames, giving the model explicit temporal information that could help it relate visual events to points in time.
Supported inputs and outputs
Seed1.5-VL accepted three primary input types:
- Text instructions and questions
- Still images
- Video
Its output was text. It did not natively generate images, video, audio, or other media. This distinction matters because the model was multimodal on the input side, but its practical role was to analyze and explain media in language.
| Capability | Supported or reported status |
|---|---|
| Text input | Yes |
| Image input | Yes |
| Video input | Yes |
| Audio input | Not reported |
| Text output | Yes |
| Image, video, or audio output | No |
| Function and tool use | Supported through the Volcano Engine API |
The technical report specified a maximum sequence-length training configuration of 131,072 tokens. That value is a context-length reference from the technical materials; it should not be interpreted as a verified current production limit because the hosted service has been retired. An exact maximum output-token limit was not verified.
What Seed1.5-VL could handle
Seed1.5-VL covered a broad set of visual-language tasks:
- Visual question answering: answering questions about objects, scenes, documents, diagrams, and events shown in an image or video.
- OCR and document understanding: reading visible text and using it as part of a larger explanation or reasoning task.
- Charts and diagrams: extracting information from visual representations and describing relationships between their components.
- Visual grounding: connecting words or descriptions to particular regions, objects, or locations in an image.
- Counting and localization: finding objects and estimating their positions or quantities.
- Three-dimensional spatial understanding: reasoning about spatial relationships represented in visual input.
- Video comprehension: tracking events across frames and relating them to temporal information.
- Visual puzzles: working through image-based problems that require multiple reasoning steps.
- GUI interaction and gameplay: interpreting screens and supporting agent-oriented tasks involving interfaces or game environments.
ByteDance reported state-of-the-art results on 38 of 60 public vision-language benchmarks in its release materials, including 14 of 19 video benchmarks and 3 of 7 evaluated GUI-agent tasks. These are provider-reported benchmark claims rather than an independent assessment, and benchmark performance does not guarantee reliable behavior on every real-world image, video, or interface.
Reasoning, coding, and tool support
Reasoning was a central part of Seed1.5-VL’s positioning. The model was built for long-chain multimodal reasoning, meaning it could combine visual evidence with several intermediate steps before producing an answer. Examples include solving a visual puzzle, interpreting a diagram and applying its information to a question, or examining a screenshot to determine which interface element should be used.
Its reasoning ability was primarily visual and language-based rather than a dedicated software-engineering specialization. It could interpret code shown in screenshots or combine visual input with programming-related instructions, but the supplied research does not establish a separate coding benchmark profile or a specialized code-generation mode.
Through the Volcano Engine API, the model supported function and tool calling. This allowed an application to define operations that the model could request during a conversation. Tool support was useful for agent experiments, GUI workflows, and systems that needed to connect visual interpretation with external actions. The model itself did not become an autonomous computer operator simply because tools were available; the surrounding application still had to implement permissions, execution, and safety controls.
Strengths and limitations
Seed1.5-VL’s main strength was the breadth of its visual reasoning target. It was not limited to image captions or basic visual question answering. Its reported capabilities covered images, video, OCR, diagrams, localization, temporal understanding, GUI interaction, and gameplay. The combination of a dedicated vision encoder and an MoE language model was intended to support both detailed visual processing and complex language-based reasoning.
The model also offered a relatively broad context configuration, with the technical report documenting up to 131,072 tokens in its maximum sequence-length training setup. That could be useful for long prompts, extended visual-language conversations, or tasks involving substantial textual context. Because the service is no longer available, however, this figure cannot be used to plan a new production deployment.
ByteDance identified several weaknesses. Fine-grained visual perception could be unreliable, particularly when the task depended on subtle differences between objects or images. Object counting, complex spatial relations, image-difference recognition, combinatorial search, mazes, and sliding puzzles were also listed as difficult areas. Some forms of temporal video reasoning remained challenging, and difficult multi-step tasks could produce unsupported assumptions or incomplete answers.
These limitations are important for practical use. A model may correctly understand the broad content of a photograph while still miscounting similar objects, overlooking a small visual distinction, or making an incorrect assumption about the order of events in a video. Applications that depend on exact localization, safety-critical interpretation, or reliable planning should validate outputs rather than treating the model’s explanation as ground truth.
API availability and lifecycle
Seed1.5-VL was hosted through Volcano Engine’s Ark API. The documented interface supported multimodal chat completions, streaming responses, a thinking-control parameter, and function tools. Its API identity was doubao-1-5-thinking-vision-pro-250428, not simply the research name Seed1.5-VL.
The model is no longer a current deployment option. Volcano Engine’s lifecycle documentation listed March 31, 2026 as the end-of-service date, and ByteDance recommended migration to doubao-seed-2-0-lite-260215. The supplied research also records a deprecation date of December 19, 2025. Since the endpoint has ended, current users should not build a new integration around the old model identifier.
No verified input or output pricing was supplied for Seed1.5-VL, and no current pricing should be inferred from historical API availability. The model’s speed and cost characteristics therefore cannot be used as a present purchasing comparison. For a new project, the relevant evaluation is whether the recommended successor or another currently supported vision-language model offers the required visual accuracy, context capacity, latency, and cost.
When Seed1.5-VL was a suitable choice
When it was available, Seed1.5-VL was a reasonable fit for research and evaluation involving multimodal reasoning rather than media generation. Suitable use cases included:
- Testing visual question answering and long-form image reasoning
- Analyzing videos and asking about events or temporal relationships
- Extracting and explaining text from documents, screenshots, charts, or diagrams
- Investigating visual grounding, counting, and localization
- Prototyping GUI-agent or gameplay-oriented systems
- Comparing multimodal benchmark performance across vision-language models
It was less appropriate for native image or video creation, audio interaction, guaranteed pixel-level accuracy, difficult combinatorial planning, or any current production system requiring a supported endpoint. A faster or less expensive vision model could be preferable for simple classification or routine OCR, while a newer supported reasoning model may be preferable for production workflows that require active maintenance and a current service-level commitment.
Bottom line
Seed1.5-VL was a technically broad vision-language model focused on understanding images and video through language-based reasoning. Its 532-million-parameter vision encoder, 20-billion-active-parameter MoE language model, video processing design, and reported benchmark results made it notable within ByteDance Seed’s 2025 multimodal research portfolio. Its strongest areas were visual reasoning, document and diagram analysis, video understanding, grounding, and agent-oriented visual tasks.
Its practical status is now the decisive consideration: the Volcano Engine service ended on March 31, 2026. Seed1.5-VL remains relevant for understanding ByteDance’s multimodal model development and for historical benchmark comparisons, but new applications should use a currently supported alternative, including the migration target identified by ByteDance, rather than the retired API.

