What is gpt-oss-safeguard-120b?
gpt-oss-safeguard-120b is an open-weight safety reasoning model released by OpenAI on October 29, 2025. It belongs to the gpt-oss family but is fine-tuned specifically for safety classification and trust and safety work rather than ordinary end-user conversation.
The model receives two important inputs: a safety policy written by the developer and the content that should be evaluated. The content can be a user message, a model completion, or a full conversation. The model then determines how the content fits the supplied policy and provides a conclusion together with reasoning that explains the decision.
This makes it different from a fixed moderation labeler whose categories and definitions are built into training. With gpt-oss-safeguard-120b, an organization can define policies for its own product, community, region, or risk area. Policies can also be revised as new risks emerge, without retraining the model for every change.
Where it fits in OpenAI's model lineup
gpt-oss-safeguard-120b is a safety-focused member of OpenAI's open-weight gpt-oss family. It is based on the gpt-oss-120b architecture and is intended to act as a moderation or safety-analysis component around other applications, including systems that use large language models.
Its role is therefore narrower than that of a general-purpose assistant. It is not presented as the main conversational model for a consumer application. It is better understood as a policy reasoning layer that can inspect inputs and outputs before, after, or alongside an application's primary model.
OpenAI distributes the weights through Hugging Face under the Apache 2.0 license. The model is not offered as a standard paid hosted model in the OpenAI API or ChatGPT according to the supplied documentation. Developers must operate it themselves or use compatible third-party infrastructure.
How policy-based classification works
A developer can provide a policy describing what should be allowed, disallowed, escalated, or otherwise labeled. The model evaluates the target content against that policy instead of relying only on a universal, provider-defined moderation taxonomy.
For example, a platform could provide rules covering harassment, self-harm content, privacy violations, or a product-specific definition of disallowed material. The model can then review a submitted message or generated answer using those rules. The same underlying checkpoint can support different policy sets for different products, provided each deployment evaluates the policies carefully.
The returned reasoning can help safety teams inspect why a decision was made, identify ambiguous policy language, and debug false positives or false negatives. The reasoning is primarily useful to developers and safety practitioners. It should not automatically be shown to ordinary users as if it were a definitive explanation of a moderation decision.
Technical specifications
| Specification | Details |
|---|---|
| Provider | OpenAI |
| Release date | October 29, 2025 |
| Model type | Open-weight safety reasoning model |
| Total parameters | Approximately 117 billion |
| Active parameters | Approximately 5.1 billion per token |
| Architecture | Mixture of experts |
| Context length | Up to 131,072 tokens, based on the underlying gpt-oss architecture |
| Input and output | Text input and text output |
| Reasoning effort | Low, medium, and high |
| License | Apache 2.0 |
| Output limit | No separate maximum output-token limit has been published for this safeguard checkpoint |
The model uses the Harmony response format and is intended to be used with that format. OpenAI's published deployment guidance identifies Transformers, vLLM, SGLang, Docker Model Runner, and other compatible open-model runtimes as possible ways to run it. OpenAI states that the published configuration can fit on a single H100 GPU, although actual operating requirements depend on quantization, runtime settings, workload, and service design.
Modalities and reasoning behavior
gpt-oss-safeguard-120b is text-only. The supplied research does not identify image, audio, or video input or output capabilities. Its intended inputs are written policies and written content such as messages, completions, and conversations.
The model supports configurable reasoning effort levels: low, medium, and high. This gives operators a way to trade decision quality and analysis depth against latency and compute consumption. Higher reasoning effort may be useful for complicated or ambiguous policies, while lower effort may be more appropriate for simpler or higher-throughput checks.
Structured Outputs are listed in the supplied technical information, but this should not be confused with a broad tool-calling or agent capability. Tool use is not verified in the supplied specifications. The model's documented purpose is policy-based classification, not external action execution or general application control.
Strengths for safety workflows
- Custom policies: Developers can define the policy at inference time rather than being restricted to a fixed moderation scheme.
- Adaptability: Policy wording can be changed as a product's rules, community standards, or risk profile change.
- Reasoned decisions: The model returns reasoning that can support investigation, policy debugging, and human review.
- Open deployment: Apache 2.0 weights allow organizations to download, modify, and deploy the model on infrastructure they control, subject to the license and their own operational responsibilities.
- Multiple review modes: It can be used for online filtering, content labeling, offline dataset review, and batch analysis.
- Defense-in-depth integration: It can serve as a deeper review layer alongside faster classifiers, human escalation, rate limits, and other product controls.
Limitations and trade-offs
The main trade-off is flexibility versus efficiency. A large reasoning model can examine nuanced policies, but it is likely to be more computationally expensive and slower than a small, purpose-built moderation classifier. That matters particularly for high-volume services where every request must be checked with very low latency.
OpenAI notes that classifiers trained on large, high-quality labeled datasets may outperform policy reasoning models on complex classification tasks. This means gpt-oss-safeguard-120b should not automatically replace a well-tested domain classifier. A practical design may use a smaller classifier for routine cases and send ambiguous, difficult, or high-impact cases to this model for deeper analysis.
The model is also not a turnkey safety system. Organizations must write clear policies, test them against representative examples, monitor outcomes, handle appeals or human review, and protect the model's reasoning from inappropriate disclosure. A policy supplied at runtime can be updated easily, but poorly written or internally inconsistent policies can still produce unreliable decisions.
There is no published hosted API price for this checkpoint because it is distributed as downloadable weights rather than sold as a standard OpenAI API model. The financial cost comes from GPU infrastructure, storage, operations, monitoring, and engineering. Self-hosting can provide control and predictable deployment boundaries, but it also transfers those responsibilities to the organization.
Coding and general-purpose use
gpt-oss-safeguard-120b is not documented as a coding-specialist model. It may process code or technical text as content to classify, but the supplied research does not establish a dedicated code-generation capability, coding benchmark, or programming-focused optimization. Its suitability should therefore be judged by the quality of its safety decisions, not by expectations associated with general coding assistants.
Likewise, it is not intended to replace a general conversational model. Using it as an end-user assistant would mismatch its documented purpose and could expose internal reasoning or produce safety judgments where ordinary dialogue is required.
Best use cases
- Filtering prompts sent to a large language model.
- Reviewing model completions before they reach users.
- Applying organization-specific rules to user-generated content.
- Labeling large datasets for moderation or safety analysis.
- Investigating borderline cases that a fast classifier cannot resolve confidently.
- Testing and refining safety policies for emerging or specialized risks.
- Running moderation workflows locally or in controlled infrastructure when a hosted endpoint is unsuitable.
When to choose this model
Choose gpt-oss-safeguard-120b when policy customization, open-weight deployment, and inspectable reasoning are more important than minimum latency or minimum infrastructure cost. It is especially relevant when an organization needs rules that are more specific than a standard moderation taxonomy, or when safety teams want to review how a model applied those rules.
A smaller dedicated classifier may be a better choice for routine, high-volume moderation where speed and cost dominate. A hosted moderation service may be more appropriate when a team does not want to operate GPUs, manage model serving, or maintain a safety evaluation pipeline. A general-purpose language model is a better fit for user-facing conversation, content generation, or coding assistance.
For many production systems, the most suitable approach is not an exclusive choice. A fast first-pass filter can handle obvious cases, while gpt-oss-safeguard-120b reviews ambiguous content, difficult conversations, policy changes, or cases that warrant human escalation.
Availability and deployment
gpt-oss-safeguard-120b is available as downloadable open weights from OpenAI's Hugging Face repository. It is described as a research preview, so teams should validate its behavior on their own policies and data before relying on it for consequential decisions.
The model works best as one layer in a broader defense-in-depth system. Deployment should include policy versioning, representative evaluations, monitoring for distribution changes, escalation paths, and controls that limit the impact of an incorrect classification. Its open-weight status provides control over deployment, but it does not remove the need for careful safety engineering.

