Seed-OSS

Seed-OSS-36B-Instruct

by ByteDance Seed · Current open-weight model; downloadable under the Apache-2.0 license

Seed-OSS-36B-Instruct is ByteDance Seed’s 36-billion-parameter Apache-2.0 open-weight language model. It supports up to 512K tokens of context, configurable reasoning budgets, streaming, and tool calling through compatible inference stacks. The text-only model is aimed at self-hosted reasoning, coding, document analysis, summarization, and agent workflows, with no verified hosted API price or fixed maximum output limit.

Text Reasoning Coding
Released on August 20, 2025, Seed-OSS-36B-Instruct is an open-weight language model for users who want to run a capable reasoning and coding model through their own infrastructure or compatible inference platforms. Its defining practical feature is a 512K-token context window, combined with configurable reasoning budgets and documented tool-calling support. The model is available from ByteDance Seed under the Apache-2.0 license and can be deployed with software such as Transformers, vLLM, and SGLang.
Outputs

What Seed-OSS-36B-Instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
5/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Seed-OSS
Model type Reasoning
Context window 524K tokens
Knowledge cutoff 07/2024
Release date 2025-08-20
Status Current open-weight model; downloadable under the Apache-2.0 license
Knowledge cutoff notes

The official Seed-OSS model card states that the training data has a knowledge cutoff of July 2024. This is the underlying training-data cutoff and is not changed by external retrieval or user-supplied context.

Model notes

Seed-OSS-36B-Instruct is a 36B-parameter dense causal language model released by ByteDance Seed on August 20, 2025. It uses RoPE, grouped-query attention, RMSNorm, and SwiGLU, with 64 layers, 80/8/8 QKV heads, a 5,120 hidden size, and a 155K vocabulary. The model is optimized primarily for international use cases and is released under Apache-2.0. The official materials describe adjustable thinking budgets, including unlimited default reasoning and recommended budget values such as 512, 1K, 2K, 4K, 8K, and 16K tokens. Official examples show tool calling through vLLM with the Seed OSS tool-call parser and support for streaming inference. No official hosted API pricing, fixed maximum output-token limit, JSON-mode guarantee, prompt-caching feature, or batch API was identified for this exact open-weight model. The model card reports a 07/2024 knowledge cutoff. Native text input and text output are documented; image, audio, and video modalities are not documented for this model.

Model guide

Seed-OSS-36B-Instruct: ByteDance Seed’s Open-Weight Long-Context Reasoning Model

Seed-OSS-36B-Instruct is ByteDance Seed’s Apache-2.0-licensed, 36-billion-parameter instruction-tuned causal language model. It is designed for self-hosted long-context reasoning, coding, tool use, question answering, summarization, and general text generation. The model supports a native context window of up to 524,288 tokens and adjustable thinking budgets, but it has no documented official hosted price, fixed maximum output-token limit, guaranteed JSON mode, or native image, audio, or video support.

What is Seed-OSS-36B-Instruct?

Seed-OSS-36B-Instruct is a 36-billion-parameter, instruction-tuned causal language model from ByteDance Seed. “Instruction-tuned” means it has been adapted to follow user requests rather than merely continue text. It can answer questions, write and explain code, summarize long documents, reason through multi-step problems, and participate in tool-use workflows.

The model is open-weight and released under the Apache-2.0 license. In practical terms, users can download the model weights and deploy them in compatible environments instead of relying on a single official consumer application or a required ByteDance-hosted endpoint. The official model materials document usage with Transformers, vLLM, SGLang, and related inference tooling.

Seed-OSS-36B-Instruct was released on August 20, 2025. Its model card identifies July 2024 as the training-data knowledge cutoff. That cutoff applies to the model’s learned knowledge; it can still process information supplied in a prompt or retrieved by an external application.

Where it fits in ByteDance Seed’s lineup

Seed-OSS-36B-Instruct is the open-weight language-model offering in ByteDance Seed’s broader research and model ecosystem. ByteDance Seed also publishes or supports models aimed at agentic productivity, coding, image generation, video generation, audio, real-time interaction, robotics, and other specialized applications. Those products do not make Seed-OSS-36B-Instruct a multimodal model: the supplied documentation describes this specific model as a text-in, text-out language model.

Its positioning is therefore different from a consumer assistant or a managed creative-generation service. The model is intended for developers, researchers, and organizations that need more control over deployment, inference configuration, data handling, and integration. It is especially relevant when a long prompt, document collection, codebase, or multi-step agent history must fit into one context window.

Key specifications

SpecificationVerified detail
ProviderByteDance Seed
Release dateAugust 20, 2025
Model familySeed-OSS
Parameters36 billion; dense causal language model
LicenseApache-2.0
Context windowUp to 524,288 tokens, or approximately 512K tokens
Knowledge cutoffJuly 2024
Input and outputText input and text output
Native image, audio, or video supportNot documented for this model
Official hosted pricingNot identified
Fixed maximum output lengthNot identified in the supplied official materials

The 512K context window is a maximum input-and-context capacity, not a promise that every deployment will process such prompts at the same speed or memory cost. Actual limits can depend on the inference engine, hardware, quantization, batch size, and configuration chosen by the operator.

Reasoning and adjustable thinking budgets

Seed-OSS-36B-Instruct is designed to support reasoning workflows in which the model spends additional generation effort working through a problem before producing its final response. The official materials describe an unlimited default reasoning setting and recommended thinking-budget values including 512, 1K, 2K, 4K, 8K, and 16K tokens.

A thinking budget is not the same as the final answer length. It is a control over how much internal reasoning work the deployment permits before the model responds. Smaller budgets can reduce latency and resource use for straightforward requests. Larger budgets may be more appropriate for difficult coding, mathematics, planning, or multi-step analysis tasks, although the supplied research does not provide benchmark results proving how accuracy changes at each setting.

This configurability gives operators a practical speed-versus-deliberation trade-off. A support or classification workflow may use a short budget, while a code-repair or research workflow may allow more reasoning. The ideal setting will depend on the task, hardware, prompt design, and serving configuration.

Coding, tool use, and agent workflows

Coding is one of the model’s documented target uses. Seed-OSS-36B-Instruct can generate code, explain existing code, help locate likely defects, and work through programming tasks using a compatible text-based deployment. Its long context is also useful for supplying larger files, specifications, logs, or repository excerpts, although the practical benefit depends on the serving system’s memory capacity and the quality of the surrounding application.

The official examples document tool calling through vLLM with the Seed OSS tool-call parser. This means an application can expose functions or tools to the model, allow the model to request one, execute that request in application code, and return the result for another model turn. The model does not automatically browse the web, execute arbitrary code, or access external systems by itself. Those capabilities must be implemented and controlled by the host application.

Tool use makes the model suitable for agentic workflows such as structured research, repository inspection, database lookup, or task planning. Developers should still validate tool arguments, restrict permissions, handle failed calls, and treat generated actions as untrusted until checked.

Supported modalities and output types

Seed-OSS-36B-Instruct is documented as a text model. It accepts text and produces text. There is no supplied evidence that this model natively accepts images, audio, or video, and it does not directly generate images, audio, video, music, speech, or other non-text media.

This limitation matters when choosing between Seed-OSS-36B-Instruct and a multimodal model in ByteDance Seed’s wider ecosystem. A developer needing image understanding, video generation, audio creation, or audiovisual interaction should select a model specifically documented for that modality rather than assuming that the Seed-OSS name represents the capabilities of every ByteDance Seed model.

Deployment, pricing, and cost

The model can be downloaded from its official Hugging Face repository and used with documented open-source inference stacks including Transformers, vLLM, and SGLang. This makes infrastructure choice a central part of the product experience. Users are responsible for selecting hardware, managing memory, configuring serving software, and operating the deployment unless they use a separate third-party host.

No official hosted API price was identified for Seed-OSS-36B-Instruct in the supplied research. It would therefore be misleading to give a per-token price or describe the model as having a standard ByteDance API tariff. A self-hosted deployment also does not mean zero cost: hardware, cloud compute, storage, engineering, monitoring, and electricity all contribute to the total cost.

The model’s cost profile is best understood as a trade-off. Its open license and downloadable weights can provide control and predictable infrastructure ownership, while a 36-billion-parameter model with a potentially 512K-token context can require substantially more memory and compute than a smaller, faster model. Long prompts and larger reasoning budgets may further increase latency and resource consumption.

Main strengths and limitations

Strengths

  • Very long context: The documented 524,288-token context window is useful for large documents, codebases, research material, and extended agent histories.
  • Open deployment model: Apache-2.0 licensing and downloadable weights support self-hosted and customized deployments.
  • Configurable reasoning: Operators can select different thinking-budget settings to balance deliberation against response time and compute use.
  • Developer-oriented integration: Official examples cover Transformers, vLLM, SGLang, streaming inference, and tool calling.
  • Broad text capability: The model is intended for reasoning, coding, question answering, summarization, and general text generation rather than one narrow task.

Limitations

  • No native media support: It is not documented as an image, audio, or video model.
  • Operational burden: Self-hosting requires suitable hardware, serving expertise, monitoring, and security controls.
  • Unknown hosted economics: No official price for a managed API was identified for this exact model.
  • Unknown fixed output ceiling: The supplied materials do not establish a universal maximum output-token limit.
  • Undocumented structured-output guarantees: No guaranteed JSON mode was identified. Applications needing strict schemas should add validation and retry logic.
  • Knowledge cutoff: The model’s learned knowledge ends at July 2024 unless the application supplies newer information through retrieval or user context.

When to choose Seed-OSS-36B-Instruct

Choose Seed-OSS-36B-Instruct when you need an open-weight model for long-context text work and are prepared to operate the inference environment. It is a strong candidate for self-hosted coding assistants, document analysis, research summarization, question answering over large supplied context, and tool-enabled agents. The Apache-2.0 license may also be important for organizations that want more control over deployment than a closed hosted service provides.

The model is particularly suitable when context size is more important than minimum latency. A large codebase, lengthy technical archive, or extended task history can be supplied within the advertised context capacity, subject to the limitations of the selected hardware and runtime. Adjustable thinking budgets also make it possible to use different reasoning settings for simple and difficult tasks.

Another model may be more appropriate when the priority is very low latency, minimal infrastructure cost, a turnkey hosted API, guaranteed structured output, or native image, audio, and video processing. A smaller model can be easier and cheaper to operate for routine classification, extraction, or short-answer tasks. A managed proprietary model may be preferable when an organization does not want to maintain serving infrastructure. A specialized multimodal model is the better choice for media input or output.

Practical verdict

Seed-OSS-36B-Instruct is best understood as a long-context, open-weight reasoning and coding model rather than a consumer chatbot or a complete AI platform. Its most meaningful differentiators are the 512K-token context capacity, configurable thinking budgets, Apache-2.0 licensing, and documented tool-use deployment patterns. Those advantages come with infrastructure responsibility and uncertain hosted pricing. For teams that value deployment control and extensive text context, it is a credible model to evaluate; for users seeking a simple multimodal application or the fastest inexpensive inference, another type of model may be a better fit.


Answers to Frequently Asked Questions

What are the main use cases and limitations of Seed-OSS-36B-Instruct?
The model is suitable for long-context document analysis, coding assistance, research summarization, question answering, reasoning, and tool-enabled agents. Its main limitations include the infrastructure burden of self-hosting, no identified official hosted API pricing, no documented universal maximum output length, no guaranteed JSON mode, and a July 2024 knowledge cutoff.
Does Seed-OSS-36B-Instruct support images, audio, or video?
No native image, audio, or video support is documented for Seed-OSS-36B-Instruct. It is a text-in, text-out model and does not directly generate images, audio, video, music, or speech.
What license does Seed-OSS-36B-Instruct use, and how can it be deployed?
Seed-OSS-36B-Instruct is released under the Apache-2.0 license. Its weights can be downloaded from the official Hugging Face repository and deployed with tools such as Transformers, vLLM, and SGLang, subject to the operator’s hardware and infrastructure requirements.
What is Seed-OSS-36B-Instruct?
Seed-OSS-36B-Instruct is a 36-billion-parameter, instruction-tuned causal language model from ByteDance Seed. It is an open-weight text model designed for reasoning, coding, summarization, question answering, and tool-use workflows.
How large is the context window of Seed-OSS-36B-Instruct?
Seed-OSS-36B-Instruct supports a documented context window of up to 524,288 tokens, or approximately 512K tokens. The practical limit and performance depend on the inference engine, hardware, quantization, batch size, and deployment configuration.


Sources 4
Provider

About ByteDance Seed