What is Seed Diffusion Preview?
Seed Diffusion Preview is an experimental diffusion language model from ByteDance Seed. Its main target is structured code generation rather than general-purpose conversation, image creation, or multimodal assistance. The model was released on July 31, 2025, as a research preview exploring whether diffusion-based language modeling can deliver substantially faster inference while retaining useful code-generation quality.
Most modern language models use autoregressive decoding: they generate one token, use that token to help generate the next, and continue sequentially. Seed Diffusion Preview takes a different approach. It represents text in discrete states and progressively refines those states, allowing parts of an output to be generated in parallel. This design is especially relevant to code, where the output has structure and where editing or filling several related regions can benefit from non-sequential generation.
The model should therefore be understood as a research-oriented code model, not as a fully documented replacement for a general assistant or a standard commercial coding API.
Where it fits in ByteDance Seed’s lineup
Seed Diffusion Preview belongs to ByteDance Seed’s research and foundation-model portfolio. It is separate from the organization’s consumer-facing creative products and from models aimed at broader agentic productivity or multimodal generation. ByteDance Seed’s current portfolio also includes models such as Seed2.1, Seedance, Seedream, SeedRealtime, and Seed Audio, but those products address different tasks. Seed Diffusion Preview’s distinguishing focus is the use of diffusion and parallel decoding for language generation, with code as its primary application.
The official model page identifies the release as a preview and provides a “Try Now” route, but the supplied public documentation does not establish a conventional public model identifier, production API contract, or universal access policy. Availability may therefore depend on the interface, account, region, or research environment through which it is offered.
How its diffusion approach works
In a diffusion language model, generation is framed as a sequence of refinement steps rather than a strictly left-to-right process. The model begins with partially specified or noisy discrete content and repeatedly improves it until the result becomes a coherent sequence. In Seed Diffusion Preview, ByteDance describes several techniques intended to make this process practical for code:
- Discrete-state diffusion: the model works with discrete language states while progressively refining a candidate output.
- Two-stage diffusion training: training is organized into two diffusion-related stages to support the generation process.
- Constrained-order learning: the model learns generation orders that preserve useful structure rather than treating every position as completely independent.
- On-policy learning: training incorporates outputs produced by the model’s own generation behavior.
- Block-wise parallel sampling: groups of tokens can be sampled together, reducing dependence on a strictly sequential decoding loop.
- KV caching: the inference system uses key-value caching to avoid repeatedly recomputing attention information during generation.
These are technical design details rather than guarantees that every application will experience the headline speed. Actual throughput can depend on hardware, output length, sampling settings, implementation, batching, and the interface used to access the model.
Speed and code-generation performance
Speed is the model’s clearest reported advantage. ByteDance reports a peak inference rate of 2,146 tokens per second on H20 GPUs and describes this as approximately 5.4 times the speed of comparable autoregressive models. This is a provider-reported result, not an independent evaluation supplied in the available research.
ByteDance also reports comparable performance to similarly scaled autoregressive models on several code benchmarks. The available information does not provide a complete benchmark table, benchmark-by-benchmark scores, or enough detail to establish how the model compares with every current coding model. The safest interpretation is that Seed Diffusion Preview is designed to investigate whether diffusion-based decoding can approach conventional code-model quality while reducing generation latency.
For users, the practical trade-off is straightforward: the model may be attractive when high-throughput generation matters more than having a mature, broadly documented product surface. It is less suitable when a project requires a guaranteed latency profile, stable API behavior, or clearly published quality results across a wide range of programming tasks.
Inputs, outputs, and supported capabilities
The supplied model information identifies Seed Diffusion Preview as a text-input and text-output model specialized for coding. Its output is code or other text, not images, audio, video, or speech. There is no verified evidence in the supplied research that it supports image, audio, or video input, nor that it produces non-text media.
| Capability | Verified status |
|---|---|
| Primary task | Structured code generation and code editing |
| Input type | Text |
| Output type | Text, including code |
| Multimodal input or output | Not supported according to the supplied model record |
| Reasoning | Editorial score: 4 out of 10; not a provider-published rating |
| Coding | Editorial score: 8 out of 10; not a provider-published rating |
| Tool or function calling | Not verified |
| Structured output or JSON mode | Not verified |
The coding score reflects an editorial assessment based on the model’s stated specialization and available code-generation results. It should not be read as an official ByteDance rating. Likewise, the reasoning score is a comparative estimate, not a published benchmark or formal capability tier.
Context, output limits, and API availability
Several practical specifications remain undocumented in the supplied sources. No verified context-window size, maximum output-token limit, conventional API model name, public token pricing, fine-tuning support, batch API, or production service-level guarantee is available. These omissions matter for engineering teams because they prevent reliable estimates of prompt capacity, maximum file size, cost per task, and operational scaling.
The model record lists caching as supported, but the accompanying notes clarify that this refers to key-value caching used during block-wise inference. It should not be interpreted as confirmation of a separate prompt-caching API feature or discounted cached-input pricing.
Similarly, the absence of a documented tool-use capability does not prove that no surrounding interface can provide tools. It means only that tool or function support is not verified for the model itself from the supplied material. Developers should avoid assuming that it can browse the web, execute code, call functions, or return schema-constrained JSON without documentation for the specific access surface.
Pricing and access
No public input-token or output-token price is supplied for Seed Diffusion Preview. The model is presented as an experimental research preview rather than a clearly documented, generally available commercial API. Consequently, there is no verified recurring price or usage rate to use for production cost calculations.
The official model page includes a way to try the model, but access details, account requirements, regional availability, quotas, and any associated charges are not established by the supplied research. Anyone evaluating it for a real application should confirm the terms shown by the current official interface instead of assuming that a research preview has the same pricing or availability as ByteDance Seed’s other developer offerings.
Best use cases
Seed Diffusion Preview is most relevant to situations where code-generation throughput and diffusion-language-model research are central requirements. Suitable uses include:
- Testing diffusion-based language generation against autoregressive code models.
- Researching parallel decoding and block-wise sampling.
- Evaluating high-throughput code completion or generation workflows.
- Experimenting with code editing, structured generation, and output refinement.
- Studying the relationship between inference speed and code quality on controlled workloads.
For example, a research team could use the model to compare how quickly several code variants can be generated, or investigate whether block-wise refinement is useful for editing a structured function. These are evaluation and experimentation scenarios rather than assurances that the model is ready to power an unattended production coding service.
When to choose this model
Choose Seed Diffusion Preview when you specifically want to explore a diffusion-based code model and the reported speed advantage is more important than a complete commercial feature set. It is particularly interesting for researchers and infrastructure teams measuring parallel decoding, throughput, and code-editing behavior on suitable hardware.
Another coding model may be more appropriate when you need a published context window, maximum output length, transparent token pricing, stable API identifiers, tool calling, fine-tuning, batch processing, or formal service guarantees. A conventional autoregressive coding model may also be preferable when predictable left-to-right generation, broad ecosystem support, and extensive production documentation matter more than experimental decoding speed.
For multimodal work, general assistant behavior, image or video generation, or integrated agent workflows, Seed Diffusion Preview is not the appropriate choice based on the supplied specifications. Those tasks belong to other model categories and ByteDance Seed products, while this preview remains centered on text-based code generation.
Overall assessment
Seed Diffusion Preview is notable because it treats inference speed as a core research problem rather than simply scaling a conventional autoregressive decoder. Its discrete diffusion design, parallel sampling strategy, and reported 2,146-token-per-second result make it a compelling subject for code-generation research.
At the same time, its experimental status limits how confidently it can be evaluated as a production service. Pricing, context length, output limits, API details, tools, and service guarantees are not verified in the supplied documentation. The model is therefore best viewed as a high-speed research preview: promising for studying diffusion-based code generation, but not yet a fully specified general-purpose coding platform.

