GPT-OSS-Safeguard

gpt-oss-safeguard-20b

by OpenAI · Research preview; currently available as an open-weight model

OpenAI's GPT-OSS-Safeguard 20B is a text-only open-weight reasoning model for customizable safety classification. It applies developer-provided policies to content, supports configurable reasoning effort and structured outputs, and is designed for self-hosted or third-party deployment rather than the OpenAI API.

Text Reasoning Coding
GPT-OSS-Safeguard 20B is a research-preview safety model fine-tuned from GPT-OSS-20B. Its defining feature is policy-conditioned classification: instead of depending only on a fixed moderation taxonomy, it interprets a safety policy supplied by the developer and uses that policy to evaluate text. The model supports low, medium, and high reasoning effort, structured outputs, and open-weight deployment, making it relevant to organizations that need customizable moderation under their own infrastructure and data-handling requirements.
Outputs

What gpt-oss-safeguard-20b can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

8/10 Reasoning
5/10 Coding
6/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family GPT-OSS-Safeguard
Model type Other
Context window 131K tokens
Release date 2025-10-29
Status Research preview; currently available as an open-weight model
Knowledge cutoff notes

No exact knowledge cutoff for gpt-oss-safeguard-20b was identified in the authoritative model documentation reviewed. The model's policy-conditioned classification behavior should not be interpreted as evidence of a newer knowledge cutoff.

Model notes

OpenAI describes this as a research-preview open-weight safety reasoning model fine-tuned from gpt-oss-20b. It has approximately 21B total parameters and 3.6B active parameters, and is intended for lower-latency or constrained environments. The model interprets developer-provided policies at inference time, supports low, medium, and high reasoning effort, provides reasoning traces for developer and safety-practitioner review, and supports Structured Outputs. It is text-only, trained for the Harmony response format, and should be used with that format. The weights are not served through the OpenAI API or ChatGPT. Hosting, compute, and infrastructure costs depend on the selected runtime or provider. OpenAI cautions that dedicated classifiers trained on large labeled datasets can outperform the model on complex risks and that reasoning-based classification can be compute-intensive.

Model guide

GPT-OSS-Safeguard 20B: Open-Weight Policy-Based Safety Reasoning

GPT-OSS-Safeguard 20B is an open-weight, text-only safety reasoning model from OpenAI. It applies developer-provided safety policies to text and is designed for content moderation, trust and safety labeling, LLM input and output filtering, and related classification workflows. With approximately 21 billion total parameters, 3.6 billion active parameters, a 131,072-token context length, configurable reasoning effort, structured outputs, and Apache 2.0 licensing, it is intended for self-hosted or third-party deployment rather than the OpenAI API or ChatGPT.

What is GPT-OSS-Safeguard 20B?

GPT-OSS-Safeguard 20B is an open-weight safety reasoning model provided by OpenAI. It is designed primarily for policy-based classification rather than general-purpose conversation. A developer supplies the rules or safety definitions that matter for a particular product, community, or review process, and the model evaluates text against those instructions.

This makes the model different from a conventional fixed-purpose moderation endpoint. For example, an organization could define categories for harassment, self-harm risk, privacy violations, or policy-specific misuse, then ask the model to classify user messages or generated responses using those definitions. The model can also be used for offline review and labeling, where latency is less important than applying a nuanced policy consistently.

The 20B model is the smaller member of the GPT-OSS-Safeguard family. OpenAI describes it as a research preview fine-tuned from GPT-OSS-20B. Its approximately 21 billion total parameters include approximately 3.6 billion active parameters, a distinction associated with its mixture-of-experts architecture. In practical terms, the active-parameter figure helps explain why the model is positioned for more constrained deployments than a model that must activate all of its parameters for every token.

Where it fits in OpenAI's model lineup

GPT-OSS-Safeguard 20B sits in OpenAI's open-weight GPT-OSS-Safeguard family, alongside the larger GPT-OSS-Safeguard 120B model. The 20B version is positioned for lower-latency or more resource-constrained environments, while the larger sibling is the relevant comparison when an organization can allocate more compute for safety reasoning.

It is not an OpenAI-hosted product in the usual API sense. OpenAI states that GPT-OSS-Safeguard models are not served through the OpenAI API and are not available in ChatGPT. The model weights are available through Hugging Face and can be run with compatible local or third-party inference systems. This gives implementers control over hosting and data location, but also means they must operate or procure the necessary infrastructure themselves.

How policy-conditioned classification works

The model's central design is bring-your-own-policy classification. Rather than asking only whether a piece of content matches a predefined label, the implementation provides the policy that defines the relevant risks and decision criteria. The model then reasons from that policy while evaluating the supplied text.

This approach can be useful when moderation requirements vary between products or change over time. A gaming community, an education platform, and an enterprise assistant may all have different definitions of unacceptable content. A policy-conditioned model can potentially adapt to those definitions without requiring a separate newly trained classifier for every policy. However, the result still depends on the quality and clarity of the policy, the examples supplied to the model, and the validation process used by the organization.

GPT-OSS-Safeguard 20B supports low, medium, and high reasoning effort. These settings allow a deployment to trade depth of analysis against latency and compute consumption. Lower effort may be more appropriate for routine screening, while higher effort may be reserved for ambiguous or higher-risk cases. OpenAI also documents access to reasoning traces for developer and safety-practitioner review. Raw reasoning traces should be treated as an operational debugging or auditing aid, not as a user-facing explanation that is automatically authoritative.

Technical specifications

SpecificationVerified detail
ProviderOpenAI
Model familyGPT-OSS-Safeguard
Model typeOpen-weight causal language model for safety reasoning
Total parametersApproximately 21 billion
Active parametersApproximately 3.6 billion
Context length131,072 tokens
Input and outputText input and text output
Reasoning effortLow, medium, and high
LicenseApache 2.0, subject to the GPT-OSS usage policy
Response formatTrained for the Harmony response format
AvailabilityOpen weights through Hugging Face; self-hosted or third-party inference

The published configuration specifies a maximum position-embedding length of 131,072 tokens. The supplied research does not identify a separate maximum output-token limit, so deployments should not assume one beyond the limits imposed by the selected runtime and available context window.

Supported modalities and outputs

GPT-OSS-Safeguard 20B is text-only. It accepts text and produces text, so it is appropriate for moderating written messages, prompts, model responses, conversations, and other textual material. The supplied specifications do not support treating it as an image, audio, or video moderation model.

The model supports structured outputs, which can make integration easier when a moderation system needs machine-readable fields such as a verdict, category, confidence-related metadata, or a rationale. Structured output support does not remove the need for application-side validation. Implementations should define an explicit schema, reject malformed responses, and establish escalation behavior for uncertain or high-impact decisions.

The supplied research does not verify general tool or function-calling support, streaming behavior, fine-tuning availability, caching, or batch API access. Those capabilities should not be assumed from the model's structured-output support.

Deployment, cost, and performance trade-offs

OpenAI documents GPT-OSS-Safeguard 20B as suitable for constrained environments, including GPUs with approximately 16 GB of VRAM depending on quantization, runtime configuration, context size, and workload. That figure is a deployment guideline rather than a universal hardware requirement. Larger contexts, higher reasoning effort, unquantized weights, and concurrent requests can increase memory and compute requirements.

The weights can be used with open inference stacks such as Transformers, vLLM, SGLang, Ollama, and llama.cpp, as well as other compatible runtimes. Hosting is not free merely because the model has an Apache 2.0 license: organizations still pay for GPUs or third-party hosting, storage, networking, monitoring, and engineering operations. No provider-hosted input or output price is supplied for this model because it is not offered through the OpenAI API.

Its main performance trade-off is between reasoning quality, speed, and infrastructure use. A reasoning-based classifier can examine complex or ambiguous policy cases more flexibly than a simple rules engine, but it may be too expensive or slow to process every item synchronously at very large scale. A practical architecture may route routine content through smaller or specialized classifiers and send uncertain, borderline, or high-risk cases to GPT-OSS-Safeguard 20B.

Strengths and limitations

Strengths

  • Customizable policies: Developers can define the safety categories and criteria used during inference.
  • Open-weight deployment: Organizations can run the model in their own environment or select a compatible third-party host.
  • Reasoning controls: Low, medium, and high effort settings support different latency and compute budgets.
  • Structured responses: Machine-readable output can simplify moderation pipelines and case-management systems.
  • Long context: The 131,072-token context length can accommodate substantial conversations or policy material, subject to runtime resources.
  • Relatively constrained deployment profile: The 3.6B active-parameter figure and documented approximately 16 GB GPU target can make it more practical than the 120B family member in some environments.

Limitations

  • Text only: It is not suitable for direct image, audio, or video moderation.
  • No OpenAI-hosted API access: Users must manage deployment or select a third-party inference provider.
  • Operational complexity: Self-hosting requires responsibility for hardware, scaling, security, updates, monitoring, and output validation.
  • Compute cost: Higher reasoning effort and large-scale synchronous classification can be expensive compared with lightweight classifiers.
  • Not always the strongest classifier: OpenAI cautions that dedicated classifiers trained on large, high-quality labeled datasets may outperform it on complex risks.
  • Policy quality matters: Ambiguous or incomplete safety policies can lead to inconsistent classifications, even when the model is capable of reasoning about the request.

Best use cases

GPT-OSS-Safeguard 20B is a good fit for teams that need adaptable moderation under their own deployment and data-control requirements. Suitable applications include filtering prompts before they reach another language model, filtering generated responses, labeling online content, reviewing conversations against an organization-specific policy, and performing offline trust and safety analysis.

It is especially relevant when the policy changes regularly or differs across products. For example, a company could maintain separate policies for a consumer chatbot, an internal assistant, and a regulated workflow, then use the same model family with different policy instructions and output schemas. Human review can be added for cases that are ambiguous, high impact, or outside the model's tested operating conditions.

When to choose GPT-OSS-Safeguard 20B

Choose GPT-OSS-Safeguard 20B when policy flexibility, self-hosting, and reasoning over difficult text cases matter more than having a simple managed moderation endpoint. It is also a reasonable option when data-residency or infrastructure-control requirements make an open-weight model preferable to a hosted service.

Choose a smaller dedicated classifier, rules engine, or routing layer when the workload is very large, decisions are straightforward, or ultra-low latency is essential. Choose a larger safety reasoning model such as GPT-OSS-Safeguard 120B when the additional infrastructure cost is justified by the need for more capacity or deeper reasoning. For image, audio, or video moderation, use a modality-specific system instead, because GPT-OSS-Safeguard 20B only processes text.

In all cases, treat the model as one component of a safety system rather than a complete policy program. Test it against representative data, validate structured responses, monitor false positives and false negatives, protect sensitive moderation inputs, and define a human escalation path for uncertain or consequential decisions.


Answers to Frequently Asked Questions

What are the limitations of GPT-OSS-Safeguard 20B?
GPT-OSS-Safeguard 20B is text-only, requires self-hosting or third-party inference, and can involve significant infrastructure and operational costs. Its results depend on the clarity of the supplied policy and the quality of validation. Dedicated classifiers may perform better for some complex risks, and high-impact or uncertain decisions should have a human escalation path.
What hardware is needed to run GPT-OSS-Safeguard 20B?
OpenAI documents the model as suitable for constrained deployments, including GPUs with approximately 16 GB of VRAM depending on quantization, runtime configuration, context size, and workload. Larger contexts, higher reasoning effort, unquantized weights, and concurrent requests may require additional memory and compute.
What are the main technical specifications of GPT-OSS-Safeguard 20B?
The model has approximately 21 billion total parameters and approximately 3.6 billion active parameters, with a context length of 131,072 tokens. It accepts and produces text, supports low, medium, and high reasoning effort, provides structured outputs, and is licensed under Apache 2.0 subject to the GPT-OSS usage policy.
What is GPT-OSS-Safeguard 20B used for?
GPT-OSS-Safeguard 20B is an open-weight safety reasoning model designed for policy-based classification. Organizations can provide their own safety rules to evaluate prompts, generated responses, user messages, conversations, and other text for risks such as harassment, self-harm, privacy violations, or policy-specific misuse.
Is GPT-OSS-Safeguard 20B available through the OpenAI API or ChatGPT?
No. GPT-OSS-Safeguard 20B is not served through the OpenAI API and is not available in ChatGPT. Its weights are available through Hugging Face for self-hosted or compatible third-party inference deployments.


Sources 5
Provider

About OpenAI