gpt-oss-safeguard

gpt-oss-safeguard-120b

by OpenAI · Research preview; open-weight and downloadable

OpenAI's gpt-oss-safeguard-120b is an Apache 2.0 open-weight reasoning model for applying developer-defined safety policies to messages, completions, and conversations. It supports explainable classification and local deployment, but requires substantial infrastructure and may be slower than dedicated moderation classifiers.

Text Reasoning Coding
gpt-oss-safeguard-120b is a research-preview safety model from OpenAI that applies a policy supplied at inference time to the content being reviewed. This design lets organizations define or update their own safety rules without retraining a classifier. The model can analyze user inputs, assistant responses, or complete conversations and return a classification conclusion with reasoning. It is available as downloadable weights under the Apache 2.0 license, so deployment, infrastructure, evaluation, and policy design remain the developer's responsibility.
Outputs

What gpt-oss-safeguard-120b can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

8/10 Reasoning
5/10 Coding
3/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family gpt-oss-safeguard
Model type Safety
Context window 131K tokens
Knowledge cutoff June 2024
Release date 2025-10-29
Status Research preview; open-weight and downloadable
Knowledge cutoff notes

The June 2024 cutoff is documented for the underlying gpt-oss pretraining data. OpenAI has not published a separate knowledge-cutoff date specifically for the gpt-oss-safeguard post-training checkpoint.

Model notes

OpenAI released this model as a research preview on October 29, 2025. It is a fine-tuned open-weight safety reasoning model based on gpt-oss-120b and is distributed under the Apache 2.0 license. The model accepts a developer-defined policy together with content to classify and returns a conclusion with reasoning. It supports low, medium, and high reasoning effort and Structured Outputs. The model is text-only, uses the Harmony response format, and is intended for safety workflows rather than general end-user interaction. It is not offered as a standard hosted model through the OpenAI API or ChatGPT. The published context length is based on the underlying gpt-oss architecture; OpenAI has not published a separate maximum output-token limit for this exact safeguard checkpoint. Pricing is not applicable to the downloadable weights, although infrastructure and inference costs still apply.

Model guide

gpt-oss-safeguard-120b: Open-Weight Model for Custom Safety Policies

gpt-oss-safeguard-120b is an OpenAI open-weight reasoning model for classifying messages, model outputs, and conversations against developer-defined safety policies. It is designed for trust and safety workflows, including customizable moderation, content labeling, and offline review, rather than general-purpose chat.

What is gpt-oss-safeguard-120b?

gpt-oss-safeguard-120b is an open-weight safety reasoning model released by OpenAI on October 29, 2025. It belongs to the gpt-oss family but is fine-tuned specifically for safety classification and trust and safety work rather than ordinary end-user conversation.

The model receives two important inputs: a safety policy written by the developer and the content that should be evaluated. The content can be a user message, a model completion, or a full conversation. The model then determines how the content fits the supplied policy and provides a conclusion together with reasoning that explains the decision.

This makes it different from a fixed moderation labeler whose categories and definitions are built into training. With gpt-oss-safeguard-120b, an organization can define policies for its own product, community, region, or risk area. Policies can also be revised as new risks emerge, without retraining the model for every change.

Where it fits in OpenAI's model lineup

gpt-oss-safeguard-120b is a safety-focused member of OpenAI's open-weight gpt-oss family. It is based on the gpt-oss-120b architecture and is intended to act as a moderation or safety-analysis component around other applications, including systems that use large language models.

Its role is therefore narrower than that of a general-purpose assistant. It is not presented as the main conversational model for a consumer application. It is better understood as a policy reasoning layer that can inspect inputs and outputs before, after, or alongside an application's primary model.

OpenAI distributes the weights through Hugging Face under the Apache 2.0 license. The model is not offered as a standard paid hosted model in the OpenAI API or ChatGPT according to the supplied documentation. Developers must operate it themselves or use compatible third-party infrastructure.

How policy-based classification works

A developer can provide a policy describing what should be allowed, disallowed, escalated, or otherwise labeled. The model evaluates the target content against that policy instead of relying only on a universal, provider-defined moderation taxonomy.

For example, a platform could provide rules covering harassment, self-harm content, privacy violations, or a product-specific definition of disallowed material. The model can then review a submitted message or generated answer using those rules. The same underlying checkpoint can support different policy sets for different products, provided each deployment evaluates the policies carefully.

The returned reasoning can help safety teams inspect why a decision was made, identify ambiguous policy language, and debug false positives or false negatives. The reasoning is primarily useful to developers and safety practitioners. It should not automatically be shown to ordinary users as if it were a definitive explanation of a moderation decision.

Technical specifications

SpecificationDetails
ProviderOpenAI
Release dateOctober 29, 2025
Model typeOpen-weight safety reasoning model
Total parametersApproximately 117 billion
Active parametersApproximately 5.1 billion per token
ArchitectureMixture of experts
Context lengthUp to 131,072 tokens, based on the underlying gpt-oss architecture
Input and outputText input and text output
Reasoning effortLow, medium, and high
LicenseApache 2.0
Output limitNo separate maximum output-token limit has been published for this safeguard checkpoint

The model uses the Harmony response format and is intended to be used with that format. OpenAI's published deployment guidance identifies Transformers, vLLM, SGLang, Docker Model Runner, and other compatible open-model runtimes as possible ways to run it. OpenAI states that the published configuration can fit on a single H100 GPU, although actual operating requirements depend on quantization, runtime settings, workload, and service design.

Modalities and reasoning behavior

gpt-oss-safeguard-120b is text-only. The supplied research does not identify image, audio, or video input or output capabilities. Its intended inputs are written policies and written content such as messages, completions, and conversations.

The model supports configurable reasoning effort levels: low, medium, and high. This gives operators a way to trade decision quality and analysis depth against latency and compute consumption. Higher reasoning effort may be useful for complicated or ambiguous policies, while lower effort may be more appropriate for simpler or higher-throughput checks.

Structured Outputs are listed in the supplied technical information, but this should not be confused with a broad tool-calling or agent capability. Tool use is not verified in the supplied specifications. The model's documented purpose is policy-based classification, not external action execution or general application control.

Strengths for safety workflows

  • Custom policies: Developers can define the policy at inference time rather than being restricted to a fixed moderation scheme.
  • Adaptability: Policy wording can be changed as a product's rules, community standards, or risk profile change.
  • Reasoned decisions: The model returns reasoning that can support investigation, policy debugging, and human review.
  • Open deployment: Apache 2.0 weights allow organizations to download, modify, and deploy the model on infrastructure they control, subject to the license and their own operational responsibilities.
  • Multiple review modes: It can be used for online filtering, content labeling, offline dataset review, and batch analysis.
  • Defense-in-depth integration: It can serve as a deeper review layer alongside faster classifiers, human escalation, rate limits, and other product controls.

Limitations and trade-offs

The main trade-off is flexibility versus efficiency. A large reasoning model can examine nuanced policies, but it is likely to be more computationally expensive and slower than a small, purpose-built moderation classifier. That matters particularly for high-volume services where every request must be checked with very low latency.

OpenAI notes that classifiers trained on large, high-quality labeled datasets may outperform policy reasoning models on complex classification tasks. This means gpt-oss-safeguard-120b should not automatically replace a well-tested domain classifier. A practical design may use a smaller classifier for routine cases and send ambiguous, difficult, or high-impact cases to this model for deeper analysis.

The model is also not a turnkey safety system. Organizations must write clear policies, test them against representative examples, monitor outcomes, handle appeals or human review, and protect the model's reasoning from inappropriate disclosure. A policy supplied at runtime can be updated easily, but poorly written or internally inconsistent policies can still produce unreliable decisions.

There is no published hosted API price for this checkpoint because it is distributed as downloadable weights rather than sold as a standard OpenAI API model. The financial cost comes from GPU infrastructure, storage, operations, monitoring, and engineering. Self-hosting can provide control and predictable deployment boundaries, but it also transfers those responsibilities to the organization.

Coding and general-purpose use

gpt-oss-safeguard-120b is not documented as a coding-specialist model. It may process code or technical text as content to classify, but the supplied research does not establish a dedicated code-generation capability, coding benchmark, or programming-focused optimization. Its suitability should therefore be judged by the quality of its safety decisions, not by expectations associated with general coding assistants.

Likewise, it is not intended to replace a general conversational model. Using it as an end-user assistant would mismatch its documented purpose and could expose internal reasoning or produce safety judgments where ordinary dialogue is required.

Best use cases

  • Filtering prompts sent to a large language model.
  • Reviewing model completions before they reach users.
  • Applying organization-specific rules to user-generated content.
  • Labeling large datasets for moderation or safety analysis.
  • Investigating borderline cases that a fast classifier cannot resolve confidently.
  • Testing and refining safety policies for emerging or specialized risks.
  • Running moderation workflows locally or in controlled infrastructure when a hosted endpoint is unsuitable.

When to choose this model

Choose gpt-oss-safeguard-120b when policy customization, open-weight deployment, and inspectable reasoning are more important than minimum latency or minimum infrastructure cost. It is especially relevant when an organization needs rules that are more specific than a standard moderation taxonomy, or when safety teams want to review how a model applied those rules.

A smaller dedicated classifier may be a better choice for routine, high-volume moderation where speed and cost dominate. A hosted moderation service may be more appropriate when a team does not want to operate GPUs, manage model serving, or maintain a safety evaluation pipeline. A general-purpose language model is a better fit for user-facing conversation, content generation, or coding assistance.

For many production systems, the most suitable approach is not an exclusive choice. A fast first-pass filter can handle obvious cases, while gpt-oss-safeguard-120b reviews ambiguous content, difficult conversations, policy changes, or cases that warrant human escalation.

Availability and deployment

gpt-oss-safeguard-120b is available as downloadable open weights from OpenAI's Hugging Face repository. It is described as a research preview, so teams should validate its behavior on their own policies and data before relying on it for consequential decisions.

The model works best as one layer in a broader defense-in-depth system. Deployment should include policy versioning, representative evaluations, monitoring for distribution changes, escalation paths, and controls that limit the impact of an incorrect classification. Its open-weight status provides control over deployment, but it does not remove the need for careful safety engineering.


Answers to Frequently Asked Questions

When should a company choose gpt-oss-safeguard-120b?
A company should consider it when custom policies, open-weight deployment, and inspectable reasoning are more important than minimum latency or infrastructure cost. A practical production design may combine it with a faster classifier, using gpt-oss-safeguard-120b for ambiguous, complex, or high-impact cases.
What are the main limitations of gpt-oss-safeguard-120b?
The model is likely to be slower and more computationally expensive than smaller moderation classifiers. Its results also depend on the clarity and consistency of the supplied policies, so organizations must perform evaluations, monitor outcomes, manage human review, and avoid treating its reasoning as an automatically definitive explanation for users.
How does gpt-oss-safeguard-120b differ from a standard moderation classifier?
Instead of relying only on fixed moderation categories, gpt-oss-safeguard-120b evaluates content against a policy supplied at inference time. Organizations can customize and update policies for their products, communities, regions, or specific risk areas without retraining the model.
How can organizations deploy gpt-oss-safeguard-120b?
The model is available as downloadable open weights through Hugging Face under the Apache 2.0 license. Organizations can run it on their own infrastructure or use compatible third-party services and runtimes such as Transformers, vLLM, SGLang, or Docker Model Runner.
What is gpt-oss-safeguard-120b used for?
gpt-oss-safeguard-120b is an open-weight safety reasoning model designed for policy-based content classification. It can evaluate user messages, model outputs, and conversations against developer-defined safety policies for moderation, dataset review, escalation, and safety analysis.


Sources 5
Provider

About OpenAI