Seed-OSS

Seed-OSS-36B-Base

by ByteDance Seed · Available open-weight model

Seed-OSS-36B-Base is ByteDance Seed’s 36-billion-parameter Apache-2.0 foundation model for self-hosted text generation. It supports up to 512K tokens of context, reasoning, mathematics, coding, summarization, and long-context research, but has no verified native multimodal, tool-calling, or JSON-mode capabilities. Its open weights provide deployment flexibility while requiring substantial infrastructure.

Text Reasoning Coding
Released on August 20, 2025, Seed-OSS-36B-Base is the synthetic-instruction-data foundation checkpoint in ByteDance Seed’s Seed-OSS family. It accepts text and produces text, and can be deployed with Transformers, vLLM, SGLang, or compatible inference systems. Unlike a hosted assistant, it is primarily a downloadable model for users who can provide their own computing infrastructure.
Outputs

What Seed-OSS-36B-Base can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Seed-OSS
Model type General Purpose
Context window 524K tokens
Release date 2025-08-20
Status Available open-weight model
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was identified in the official model card, repository, or configuration documentation.

Model notes

Seed-OSS-36B-Base is the synthetic-instruction-data foundation checkpoint in the Seed-OSS family. ByteDance also released Seed-OSS-36B-Base-woSyn without synthetic instruction data and Seed-OSS-36B-Instruct as the post-trained variant. The model uses a causal language architecture with GQA, 64 layers, 80 attention heads, 8 key-value heads, 5,120 hidden size, 155,136 vocabulary size, BF16 weights, and a 512K maximum context. The model card describes text input and text output and does not document native image, audio, or video capabilities. Official deployment examples use Transformers, vLLM, and SGLang. Editorial scores are comparative estimates rather than vendor-provided ratings.

Cost

Model pricing

Input No official hosted API price; downloadable weights are available under Apache-2.0
Output No official hosted API price; deployment cost depends on user infrastructure or third-party hosting
Model guide

Seed-OSS-36B-Base: ByteDance’s Open-Weight Foundation Model for 512K-Token Context

Seed-OSS-36B-Base is ByteDance Seed’s 36-billion-parameter, Apache-2.0 open-weight causal language model. It is designed for self-hosted research and development, with native support for up to 512K tokens of text context, flexible thinking-budget control, and broad general-language, reasoning, mathematics, coding, and agent-oriented capabilities.

What is Seed-OSS-36B-Base?

Seed-OSS-36B-Base is a 36-billion-parameter causal language model released by ByteDance Seed. It is distributed as open weights under the Apache-2.0 license, making it suitable for research, self-hosted experiments, application development, and model customization subject to the license and the model’s usage guidance.

In practical terms, the model predicts and generates text. It can be used for general language modeling, summarization, question answering, coding assistance, reasoning experiments, long-document processing, and other text-based workflows. It is not a consumer chatbot with a unified web interface, and the supplied research does not identify an official hosted API price for this checkpoint.

The model was released on August 20, 2025, alongside Seed-OSS-36B-Base-woSyn and Seed-OSS-36B-Instruct. The selected Base version includes synthetic instruction data during pretraining. The woSyn checkpoint omits that data, while the Instruct version is the post-trained variant intended to follow user instructions more directly.

Where it fits in the Seed-OSS family

Seed-OSS-36B-Base is best understood as a foundation checkpoint rather than a ready-made assistant. Foundation models provide the underlying language and reasoning capabilities from which researchers or developers can build other systems. They may require more careful prompting, fine-tuning, or additional post-training than an instruction-tuned model.

That distinction matters when choosing between the family’s checkpoints. Seed-OSS-36B-Base is the relevant option for examining or adapting the base model and for experiments that benefit from the synthetic-instruction-data training recipe. Seed-OSS-36B-Base-woSyn is positioned for research comparing the effect of that data. Seed-OSS-36B-Instruct is generally the more appropriate sibling when the immediate goal is assistant-style instruction following or documented tool-oriented deployment.

Architecture and 512K-token context limit

The verified configuration describes a model with 64 layers, 80 attention heads, and 8 key-value heads. The 8 key-value heads implement grouped-query attention, or GQA, a design that can reduce key-value memory requirements compared with using a separate key-value head for every attention head. The model also uses RoPE positional embeddings, RMSNorm, and SwiGLU activation.

Its hidden size is 5,120, and its vocabulary contains 155,136 tokens. The configuration specifies a maximum position embedding length of 524,288 tokens, equivalent to 512K tokens. ByteDance describes the model as natively supporting this context length. This is useful for long documents, large code repositories, extended transcripts, and research workflows that need to keep substantial material in one prompt.

A 512K context window is a maximum input-context capability, not a guarantee that every deployment will handle it economically or quickly. Actual throughput and memory requirements depend on the precision, serving framework, hardware, batching, prompt length, and generation settings. The supplied specifications do not identify a maximum output-token limit, so that value should be treated as unknown rather than assumed to equal the context limit.

Capabilities and supported modalities

Seed-OSS-36B-Base is a text-input, text-output model. It does not provide native image, audio, or video input or output according to the supplied model information. Its main capability areas are general language generation, long-context processing, reasoning, mathematics, coding, summarization, question answering, and agent-oriented text workflows.

The release emphasizes flexible thinking-budget control. This allows a deployment or prompt strategy to balance how much internal reasoning effort is used against response latency and resource consumption. The available research supports describing this as a model feature, but it does not provide a universal quality threshold, benchmark table, or guaranteed performance level for a particular thinking budget.

For coding, the model can serve as a self-hosted text model for code explanation, generation, transformation, debugging discussions, and repository-level analysis when the relevant files fit within the deployment’s usable context. It should not be confused with a complete coding agent: the model itself does not supply a documented execution environment, browser, filesystem, or external service access.

Tools, structured output, and serving

The supplied model record does not verify native function or tool-calling support for Seed-OSS-36B-Base, so tool use is rated as unsupported or undocumented for this specific checkpoint. The Seed-OSS release documentation provides clearer tool-calling deployment instructions for Seed-OSS-36B-Instruct, but those instructions should not automatically be transferred to the Base model.

Similarly, no distinct native JSON mode or guaranteed structured-output feature is verified for this model. A developer may be able to constrain or validate generated text in an external serving stack, but that is different from a provider-guaranteed structured-output capability.

Official deployment paths include Transformers, vLLM, and SGLang. The weights are provided in BF16 safetensors format. Because this is a 36-billion-parameter model, full-precision or BF16 serving requires substantial hardware resources. Quantization and tensor parallelism may be necessary for practical local deployment, particularly when using long contexts or serving multiple requests.

The model record lists streaming support, but streaming behavior can depend on the selected inference framework and integration. It also lists fine-tuning as supported. The research does not specify a single official fine-tuning recipe, hardware requirement, or maximum concurrent-request configuration.

Pricing and operating cost

There is no official hosted API price supplied for Seed-OSS-36B-Base. The weights are downloadable under the Apache-2.0 license, but downloading a model is not the same as running it at no cost. Users must account for GPU or accelerator hardware, storage, electricity, engineering time, and any third-party hosting charges.

This cost structure makes the model different from a metered commercial API. A self-hosted deployment may be attractive for organizations that need control over infrastructure, data handling, or model customization. It may be less economical for occasional users who would otherwise pay only for a small number of API requests. Long 512K-token prompts can also increase memory use and reduce throughput, so the maximum context should not be treated as a promise of low-cost inference.

The supplied editorial assessment gives the model a comparative cost score of 8 out of 10 and a speed score of 5 out of 10. These are editorial estimates, not ByteDance-published measurements. They reflect the model’s open-weight availability and potentially favorable software cost alongside the heavier infrastructure and latency trade-offs associated with a 36B model and very long contexts.

Main strengths and limitations

Strengths

  • Very long context: The 524,288-token configuration supports documents and code collections that are far beyond the context size of many conventional deployments.
  • Broad text capability: The model targets general language, reasoning, mathematics, coding, summarization, and agent-oriented research rather than a narrow single task.
  • Open-weight deployment: Users can download and serve the BF16 weights through established open-source inference tooling instead of depending on one hosted endpoint.
  • Research flexibility: Its foundation-model status and relationship to the Base-woSyn and Instruct checkpoints make it useful for studying training and post-training choices.
  • Adjustable reasoning effort: Flexible thinking-budget control can help deployments make a deliberate trade-off between reasoning effort, latency, and resource consumption.

Limitations

  • Not an instruction-tuned assistant: It may need careful prompts, fine-tuning, or additional post-training to behave consistently in assistant workflows.
  • Heavy infrastructure requirements: A 36B-parameter BF16 model with a large context window can require substantial memory and engineering effort.
  • No verified native multimodality: It handles text only and is not suitable for direct image, audio, or video tasks.
  • Unverified tool and JSON guarantees: The supplied specifications do not confirm native function calling or a dedicated JSON mode for this Base checkpoint.
  • No known output-token maximum: The context configuration is documented, but the research does not establish a separate maximum generated-output limit.
  • Safety and reliability boundaries: The model card cautions against professional medical, legal, or financial advice and high-impact automated decisions without rigorous evaluation and human oversight.

When to choose Seed-OSS-36B-Base

Choose Seed-OSS-36B-Base when you need an open-weight text model for self-hosted research or development and have the infrastructure to serve a relatively large checkpoint. It is especially relevant for long-context experiments, code and document analysis, reasoning research, fine-tuning investigations, and applications where keeping model execution under your own operational control is important.

It is also a reasonable choice when the 512K-token context is central to the workload. Examples include analyzing a large technical document set, examining an extensive codebase, or testing how a foundation model handles long sequences. The model’s open license and compatibility with Transformers, vLLM, and SGLang can simplify integration for teams already operating open-model infrastructure.

Choose another option when you need a turnkey chatbot, a predictable hosted API bill, low-latency responses on modest hardware, or built-in multimodal input and output. A smaller model may be preferable when speed and serving cost matter more than maximum context or model capacity. Seed-OSS-36B-Instruct is likely the better fit within the same family when direct instruction following and documented assistant or tool-use behavior are more important than working with the base checkpoint.

Bottom line

Seed-OSS-36B-Base is a substantial open-weight foundation model aimed at users who want control over deployment rather than a ready-made online assistant. Its defining practical features are the 36B scale, Apache-2.0 distribution, text-only design, native 512K-token context, and compatibility with common self-hosting frameworks. Those advantages come with significant compute requirements and the need to handle instruction following, structured responses, tool integration, and safety evaluation at the application level.


Answers to Frequently Asked Questions

How can Seed-OSS-36B-Base be deployed and what does it cost?
The model can be deployed with Transformers, vLLM, or SGLang using BF16 safetensors weights. It has no supplied official hosted API price because it is distributed as downloadable open weights. Operating costs still include suitable GPU or accelerator hardware, storage, electricity, engineering, and possible third-party hosting fees. Quantization and tensor parallelism may be needed for practical deployment.
Can Seed-OSS-36B-Base be used for tool calling, JSON output, or multimodal tasks?
Seed-OSS-36B-Base is designed for text input and text output and does not have verified native image, audio, or video capabilities. Native function calling, tool use, and a guaranteed JSON mode are also not confirmed for this Base checkpoint. External serving tools may constrain or validate outputs, but those features are not guaranteed by the model itself.
What is Seed-OSS-36B-Base?
Seed-OSS-36B-Base is a 36-billion-parameter causal language model released by ByteDance Seed as open weights under the Apache-2.0 license. It is a text-input, text-output foundation model for research, self-hosted applications, long-context processing, coding, reasoning, summarization, and model customization.
How large is the context window of Seed-OSS-36B-Base?
Seed-OSS-36B-Base has a documented maximum position embedding length of 524,288 tokens, equivalent to a 512K-token context window. This supports large documents, extensive code repositories, and long transcripts, although actual performance, memory use, and cost depend on the hardware, precision, serving framework, and prompt length.


Sources 4
Provider

About ByteDance Seed