Model catalog

ByteDance Seed Models

Browse the AI models associated with ByteDance Seed. Compare current and historical models by family, capabilities, context window, availability and intended use.

24 models tracked
24 Total models
17 Model families
7 Model types
22 Current / accessible
All models

ByteDance Seed model catalog

Long-horizon, high-precision dexterous robot manipulation and real-world VLA policy specialization

Type Other
Reasoning 7/10
Speed 4/10
Multimodal Image input Fine-tuning
Status

Current research model/framework; public API availability not documented

View model →
ByteDance Seed logo
Protenix

Protenix

Protein and biomolecular complex structure prediction, computational biology research, molecular design workflows, and self-hosted scientific inference

Type Other
Reasoning 8/10
Speed 5/10
Fine-tuning
Status

Active open-source biomolecular structure prediction project with multiple model variants, including Protenix-v1 and Protenix-v2

View model →

Multimodal agent workflows, search and information retrieval, coding agents, GUI interaction, image and video understanding, complex instruction following, and long-context business tasks.

Type Multimodal
Context 256K
Reasoning 8/10
Speed 8/10
Multimodal Image input Video input
Status

Beta; accessible through BytePlus ModelArk as seed-1-8-251228. The separate LAS multimodal deep-thinking operator using Seed1.8 ended service on 2026-09-20.

Input USD 0.25 per million tokens for prompts up to 128K tokens; USD 0.50 per million tokens for prompts over 128K and up to 256K; cached input USD 0.05 per million tokens.
Output USD 2.00 per million tokens for prompts up to 128K tokens; USD 4.00 per million tokens for prompts over 128K and up to 256K. Batch output pricing is USD 1.00 or USD 2.00 per million tokens by prompt-length tier.
View model →

Self-hosted general language modeling, long-context research, reasoning experiments, coding assistance, summarization, and foundation-model fine-tuning

Type General Purpose
Context 524K
Reasoning 7/10
Speed 5/10
Fine-tuning Streaming
Status

Available open-weight model

Input No official hosted API price; downloadable weights are available under Apache-2.0
Output No official hosted API price; deployment cost depends on user infrastructure or third-party hosting
View model →

Self-hosted long-context reasoning, coding, agentic tool use, research, question answering, summarization, and general text generation

Type Reasoning
Context 524K
Reasoning 8/10
Speed 5/10
Tool use Streaming
Status

Current open-weight model; downloadable under the Apache-2.0 license

View model →

Research, custom post-training, long-context processing, coding experiments, and self-hosted foundation-model deployment

Type General Purpose
Context 512K
Reasoning 7/10
Speed 4/10
Tool use Fine-tuning
Status

Current open-weight downloadable foundation model

View model →

General-purpose Chinese and multilingual assistance, coding, reasoning, image and document understanding, and voice-interaction applications.

Type Multimodal
Context 33K
Reasoning 7/10
Speed 8/10
Multimodal Image input Audio input
Status

Retired; the primary doubao-1-5-pro-32k-250115 deployment was scheduled to shut down on 2026-09-21 at 14:00 China Standard Time.

View model →
ByteDance Seed logo
Seed1.5

Seed1.5-VL

Visual reasoning, image and video understanding, OCR, visual grounding, GUI-agent research, gameplay analysis, and multimodal benchmark evaluation

Type Multimodal
Context 131K
Reasoning 8/10
Speed 6/10
Multimodal Image input Video input
Status

Retired; Volcano Engine service ended on 2026-03-31

View model →

Cross-modal semantic search, text-image retrieval, video retrieval, multimodal knowledge bases, classification, clustering, and recommendation

Type Other
Context 128K
Reasoning 1/10
Speed 8/10
Multimodal Image input Video input
Status

Current and available through Volcano Engine

View model →
ByteDance Seed logo
Seed1.6

Seed1.6

Multimodal document analysis, visual question answering, coding, mathematics, general reasoning, long-context analysis and adaptive-thinking applications

Type Multimodal
Context 256K
Reasoning 8/10
Speed 7/10
Multimodal Image input Tool use
Status

Deprecated; new endpoint creation stopped on 2026-09-24; existing service scheduled for automatic migration or replacement on 2026-11-24

View model →

Deep reasoning, coding, mathematics, logical analysis, document understanding, and visual reasoning over images or videos

Type Reasoning
Context 256K
Reasoning 8/10
Speed 5/10
Multimodal Image input Video input
Status

Deprecated; scheduled for retirement and migration to a Seed2.0 model

View model →

Cost-conscious production applications requiring long-context multimodal understanding, document and video analysis, coding assistance, tool use, GUI automation, and structured extraction.

Type Multimodal
Context 262K
Reasoning 8/10
Speed 8/10
Multimodal Image input Audio input
Status

Current; latest documented release seed-2-0-lite-260428

Input $0.25 per 1M non-audio input tokens for prompts up to 128K; $0.50 per 1M for prompts above 128K and up to 256K. Audio input: $3.75 per 1M tokens up to 128K and $7.50 per 1M above 128K.
Output $2.00 per 1M output tokens for prompts up to 128K; $4.00 per 1M for prompts above 128K and up to 256K.
View model →

High-concurrency inference, batch generation, classification, extraction, summarization, and cost-sensitive multimodal workloads

Type Lightweight
Context 256K
Reasoning 7/10
Speed 9/10
Multimodal Image input Video input
Status

Current; available through ByteDance's Volcano Engine model API

Input $0.03 per 1M input tokens
Output $0.31 per 1M output tokens
View model →

Complex multimodal reasoning, long-chain agent workflows, visual and video analysis, document understanding, scientific research support, coding, and enterprise automation.

Type Multimodal
Context 200K
Reasoning 9/10
Speed 6/10
Multimodal Image input Audio input
Status

Active and currently accessible through Volcano Engine Ark; canonical deployment version 260215

Input CNY 3.2 per million tokens for 0–32K input; CNY 4.8 per million tokens for 32–128K input; CNY 9.6 per million tokens for 128–256K input. Batch inference input pricing is CNY 1.6, 2.4, and 4.8 per million tokens for the same tiers.
Output CNY 16 per million tokens for 0–32K input; CNY 24 per million tokens for 32–128K input; CNY 48 per million tokens for 128–256K input. Batch inference output pricing is CNY 8, 12, and 24 per million tokens for the same tiers.
View model →

Complex agent workflows, high-value office and research tasks, long-horizon coding, document and visual analysis, video understanding, and tool-enabled productivity automation.

Type General Purpose
Reasoning 8/10
Speed 6/10
Multimodal Image input Video input
Status

Current and officially released; available through ByteDance Seed, Doubao, and Volcano Engine API channels.

View model →

Fast multimodal agents, coding assistants, document and video analysis, tool-calling workflows, structured data extraction, and cost-sensitive production applications

Type Multimodal
Context 256K
Reasoning 8/10
Speed 9/10
Multimodal Image input Video input
Status

Current and available through Volcano Engine Ark; API model identifier doubao-seed-2-1-turbo-260628

Input CNY 6 per 1M input tokens for input lengths up to 256K tokens
Output CNY 30 per 1M output tokens
View model →

Long-form audiovisual storytelling, text-to-video, reference-based video generation, creative production, advertising, education, industrial simulation, and video editing.

Type Multimodal
Reasoning 2/10
Speed 7/10
Multimodal Image input Audio input
Status

Current; available through ByteDance platforms including Jimeng AI, Doubao Pro, and the Seed platform. API access through BytePlus ModelArk was announced as forthcoming.

View model →
ByteDance Seed logo
Seedance 2.0

Seedance 2.0

Multimodal text-to-video and reference-based video creation, cinematic short clips, video editing and extension, multi-shot storytelling, and synchronized audio-video production.

Type Multimodal
Reasoning 1/10
Speed 6/10
Multimodal Image input Audio input
Status

Current official model page remains available; older generation superseded by Seedance 2.5

View model →
ByteDance Seed logo
Seed Audio

Seed Audio 1.0

Full-scene audio creation, expressive voice generation, dubbing, dialogue, sound effects, ambience, advertising, games, podcasts, and multilingual audio production

Type Audio Generation
Multimodal Image input Audio input
Status

Currently available through BytePlus

View model →

High-throughput code generation research, diffusion-language-model evaluation, and code-editing experiments

Type Coding
Reasoning 4/10
Speed 10/10
Status

Experimental research preview

View model →
ByteDance Seed logo
Seed GR

Seed GR-3

Embodied robotics research, long-horizon manipulation, bimanual control, dexterous object handling, and adapting robot policies to new objects and tasks

Type Other
Reasoning 6/10
Speed 6/10
Multimodal Image input Media output
Status

Current officially documented research model; no public hosted API, commercial pricing, or downloadable weights verified

View model →
ByteDance Seed logo
SeedRealtime

SeedRealtime

Real-time audio-visual assistants, scene-aware guidance, live explanation, interactive learning, accessibility, and proactive multimodal collaboration

Type Multimodal
Reasoning 7/10
Speed 9/10
Multimodal Image input Audio input
Status

Current; fully rolled out for large-scale audio-visual full-duplex deployment

View model →
ByteDance Seed logo
Seedream 5.0

Seedream 5.0 Pro

Professional image generation and editing, high-density infographics, advertising and e-commerce creative, multilingual visual content, precise spatial edits, layer separation, and multi-reference image workflows

Type Multimodal
Reasoning 8/10
Speed 7/10
Multimodal Image input Media output
Status

Current and publicly documented; available through BytePlus ModelArk and ByteDance Seed services

Input First reference image free; additional reference images $0.003 per image
Output $0.045 per output image for images up to 2.36 million pixels; $0.09 per output image above 2.36 million pixels
View model →
ByteDance Seed logo
UI-TARS-1.5

UI-TARS-1.5-7B

Open-weight computer-use research, GUI grounding, browser automation prototypes, screenshot-based interface interaction, and visual action-model experimentation

Type Multimodal
Context 128K
Reasoning 8/10
Speed 5/10
Multimodal Image input
Status

Available open-weight research release; superseded by UI-TARS-2 for the provider's newer UI-TARS development line

View model →