gpt-oss

gpt-oss-20b

by OpenAI · Current open-weight model; downloadable and usable through self-hosted or third-party inference infrastructure. Not served through the OpenAI API or ChatGPT.

OpenAI's gpt-oss-20b is a 21-billion-parameter open-weight reasoning model for local, private and specialized deployments. It offers a 128K context window, configurable reasoning effort, tool use, structured outputs and fine-tuning under an Apache 2.0 license, but it is text-only and has no official OpenAI API pricing.

Text Reasoning Coding
Released on August 5, 2025, gpt-oss-20b gives developers an OpenAI reasoning model that they can download and operate on infrastructure they control. Its mixture-of-experts design activates approximately 3.6 billion parameters per token, while the released MXFP4-quantized weights are intended to run within roughly 16 GB of memory. The model is aimed at local inference, private deployments, coding, agentic workflows and experimentation rather than turnkey multimodal applications.
Outputs

What gpt-oss-20b can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Fine-tuning Structured output
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family gpt-oss
Model type Reasoning
Context window 131K tokens
Release date 2025-08-05
Status Current open-weight model; downloadable and usable through self-hosted or third-party inference infrastructure. Not served through the OpenAI API or ChatGPT.
Knowledge cutoff notes

OpenAI's public model documentation reviewed for this record does not specify an exact knowledge cutoff for gpt-oss-20b. Web browsing or search tools can provide newer information during inference but do not establish or change the underlying cutoff.

Model notes

OpenAI released gpt-oss-20b on August 5, 2025. It has approximately 21B total parameters and 3.6B active parameters per token, uses a mixture-of-experts architecture, and supports low, medium, and high reasoning effort. The model uses OpenAI's Harmony response format. It is distributed under Apache 2.0 subject to the gpt-oss usage policy. OpenAI provides no hosted API endpoint or per-token price for this model; users are responsible for compute, storage, operations, or third-party hosting costs. The published model card describes the model as text-only. Tool use, streaming, and structured outputs depend partly on the serving runtime, although OpenAI's reference ecosystem supports these workflows. The 128K context value comes from the model configuration and OpenAI's published architecture documentation. A provider-published maximum output-token limit was not found.

Cost

Model pricing

Input No official OpenAI API input price; self-hosting and third-party hosting costs vary.
Output No official OpenAI API output price; self-hosting and third-party hosting costs vary.
Model guide

gpt-oss-20b: Features, Context, Pricing and Local Deployment

gpt-oss-20b is OpenAI's 21-billion-parameter open-weight reasoning model for local, private and specialized deployments. It provides configurable reasoning effort, tool use, structured outputs and fine-tuning support within a 128K-token context window, but it is text-only and is not available through the OpenAI API or ChatGPT.

What is gpt-oss-20b?

gpt-oss-20b is an open-weight reasoning language model provided by OpenAI and released on August 5, 2025. “Open-weight” means that the model weights can be downloaded and run by developers instead of being available only behind a managed provider endpoint. Users can deploy it locally, on-premises, in a private cloud or through a third-party inference host.

The model has approximately 21 billion total parameters. It uses a mixture-of-experts architecture, so only a subset of those parameters is active for each token. OpenAI documents approximately 3.6 billion active parameters per token. This design aims to provide a larger model's representational capacity while reducing the computation required for each generated token.

gpt-oss-20b is the smaller model in the original gpt-oss release, alongside gpt-oss-120b. The two models share the open-weight deployment approach, but gpt-oss-20b is positioned for lower-latency and more hardware-conscious deployments. It is not included in ChatGPT and is not served as an official model through the OpenAI API.

Who should use gpt-oss-20b?

The model is primarily intended for developers and organizations that need more control over inference than a hosted API normally provides. Running the weights yourself can support private processing, custom safety policies, specialized tool integrations and fine-tuning with organization-specific data.

It is especially relevant for coding assistants, internal question-answering systems, local research tools, agentic workflows and applications that need a reasoning model without sending every prompt to a provider-managed service. It can also be useful for rapid experimentation because the weights and reference integrations are available through the OpenAI model repository and related open-source tooling.

This flexibility comes with operational responsibility. A self-hosted deployment requires decisions about hardware, serving software, access control, monitoring, logging, abuse prevention, updates and evaluation. The model is therefore not a drop-in replacement for a managed API if the main priority is immediate production deployment with minimal infrastructure work.

Architecture and 128K context window

gpt-oss-20b uses a 24-layer Transformer architecture with mixture-of-experts routing. The configuration includes 32 local experts and activates four experts per token. It also uses alternating sliding and full attention patterns, grouped multi-query attention and rotary positional embeddings. These details affect how the model processes long inputs and how efficiently an inference runtime can serve it.

The documented maximum context length is 128K tokens, or 131,072 tokens in the model configuration. The context window includes the information supplied to the model during a request, such as system instructions, conversation history, documents, tool definitions and the generated response. A long context limit does not guarantee that every runtime will perform equally well at the maximum length; memory use and latency depend on context size, batching, hardware and serving software.

The released checkpoint uses MXFP4 quantization for the mixture-of-experts weights. Quantization stores model values in a more compact numerical format, reducing memory requirements compared with higher-precision weights. OpenAI states that gpt-oss-20b can run within approximately 16 GB of memory, but this is a target rather than a universal hardware requirement. Runtime overhead, operating-system memory, context length, concurrent requests and implementation details can increase the actual requirement.

Reasoning, coding and general capabilities

gpt-oss-20b is designed as a reasoning model rather than a basic text-completion model. It supports low, medium and high reasoning-effort settings. These settings let an application trade deliberation depth against response speed: lower effort can reduce latency, while higher effort may be more suitable for difficult multi-step tasks. The exact quality and speed difference depends on the prompt, runtime and hardware.

The model was trained for instruction following, STEM tasks, coding, general knowledge and agentic tool-use scenarios. In practical terms, it can be used to explain technical material, generate or review code, analyze text, follow structured instructions and break a problem into multiple steps before producing an answer.

OpenAI reports that gpt-oss-20b performs similarly to OpenAI o3-mini on several published evaluations. OpenAI's comparison information lists scores of 85.3 on MMLU, 71.5 on GPQA Diamond, 17.3 on Humanity's Last Exam, 96.0 on AIME 2024 and 98.7 on AIME 2025. These are provider-reported benchmark results under the stated evaluation conditions. They are useful for understanding the intended positioning, but they should not be treated as a guarantee of performance on a particular application or dataset.

The model's coding capability is best understood as part of its broader reasoning and instruction-following design. It can support coding assistants, code explanation, debugging and tool-connected development workflows. The supplied research does not establish a universal programming-language ranking or a guaranteed software-engineering success rate, so production use should include repository-specific tests and human review.

Tools, function calling and structured output

gpt-oss-20b was trained for tool use and can be connected to function-calling workflows. A tool-connected application can ask the model to select an available function, provide arguments and incorporate the returned result into a later response. The model itself does not automatically gain access to the internet, databases or operating-system actions; those capabilities come from the tools and runtime that the operator connects.

OpenAI's reference ecosystem includes browser-based web browsing and Python execution examples. The model is also compatible with OpenAI-compatible local servers and several inference runtimes. Tool behavior, schemas, error handling and security controls can vary between Transformers, vLLM, SGLang, Ollama, LM Studio and other serving stacks.

Web search is an external capability rather than a change to the model's underlying knowledge. If a browser or search tool is connected, it supplies current information during inference. It does not establish a new knowledge cutoff or update the model weights.

The model uses OpenAI's Harmony response format. OpenAI recommends using this format for correct behavior when reasoning messages, tool calls or structured responses are involved. gpt-oss-20b also supports Structured Outputs when the serving implementation exposes that capability. Developers should verify the exact schema-enforcement behavior of their selected runtime instead of assuming that every OpenAI-compatible server provides identical guarantees.

Input and output modalities

gpt-oss-20b is text-only. It accepts text prompts and produces text responses, reasoning-related content, tool calls and structured machine-readable responses depending on the prompt format and serving implementation.

It does not natively accept images, audio or video. It also does not generate images, audio, music or video. An application can place extracted text or externally generated descriptions into a prompt, but that is not the same as native multimodal understanding. For image analysis, speech interaction or video generation, a model designed for those modalities is more appropriate.

Deployment options and hardware

OpenAI's official repository provides reference implementations and integrations for Transformers, vLLM, SGLang, Ollama, LM Studio, PyTorch, Triton and Apple Metal. The ecosystem also includes examples for downloading weights from Hugging Face and running the model through OpenAI-compatible local servers.

The approximately 16 GB memory target makes gpt-oss-20b more practical for some consumer GPUs, workstations and selected edge systems than larger open-weight models. However, “can run” and “can serve efficiently” are different requirements. A short single-user prompt may fit within a hardware target while long contexts, multiple simultaneous users, higher throughput or a large key-value cache require substantially more memory.

Runtime selection also affects speed. A well-optimized implementation with supported quantization and suitable hardware may deliver responsive local inference, while unsupported formats or constrained memory can cause slower generation, CPU offloading or reduced concurrency. Benchmarking the intended workload is more reliable than relying only on the nominal parameter count.

Pricing, license and API availability

There is no official OpenAI per-token input or output price for gpt-oss-20b because OpenAI does not host this model through the OpenAI API. The weights are available for download under the Apache 2.0 license, subject to OpenAI's gpt-oss usage policy.

“Free to download” does not mean that deployment has no cost. Operators may pay for a GPU or workstation, cloud instances, storage, electricity, networking, maintenance and engineering time. Third-party providers may offer hosted inference, but their prices depend on the provider, hardware, request volume, performance target and deployment configuration. Those prices should not be confused with an official OpenAI API rate.

The Apache 2.0 license generally supports commercial and private use subject to its terms, while the separate usage policy and applicable laws still matter. Organizations should review both before deploying the model in a product or regulated workflow.

Main strengths and limitations

Strengths

  • Local control: The weights can be run on infrastructure managed by the developer or organization.
  • Reasoning controls: Low, medium and high reasoning effort provide a practical latency-versus-deliberation trade-off.
  • Long context: The documented 128K-token context supports lengthy instructions, documents and tool histories, subject to runtime memory limits.
  • Efficient architecture: Mixture-of-experts routing activates approximately 3.6 billion parameters per token rather than all approximately 21 billion parameters.
  • Customization: Developers can adjust prompts, tools, safety controls, runtimes and fine-tuning data.
  • Open deployment ecosystem: Official references cover several widely used local and self-hosted runtimes.

Limitations

  • No managed OpenAI endpoint: It is not available through the OpenAI API or ChatGPT.
  • Text-only operation: Native image, audio and video input and output are not supported.
  • Operational burden: The operator is responsible for scaling, security, monitoring, tool execution and safety evaluation.
  • Variable runtime behavior: Streaming, tool use and structured outputs depend partly on the selected serving implementation.
  • No documented maximum output limit in the supplied research: A provider-published maximum output-token value was not identified.
  • Potentially variable quality: Benchmark results do not guarantee performance on a specific domain, codebase or business process.

When to choose gpt-oss-20b

Choose gpt-oss-20b when local or private inference is more important than access to a fully managed API. It is a sensible candidate for a coding assistant running near the development environment, an internal document assistant, a private agent with organization-controlled tools or an edge deployment where a larger model would be too slow or expensive.

It can also be attractive when an application needs to experiment with reasoning effort, function calling, structured responses or fine-tuning without committing to a provider-hosted model. Its smaller active parameter count and memory target may offer a better speed-and-cost profile than larger open-weight models, especially for single-user or moderate-throughput deployments.

Another option may be more appropriate when the application needs native image, audio or video capabilities, guaranteed managed availability, a documented hosted price, or minimal infrastructure ownership. A larger reasoning model may be preferable for tasks where maximum quality is more important than local resource requirements. Conversely, a smaller non-reasoning model may be preferable for simple classification or short responses where deliberation adds unnecessary latency.

For a deployment decision, compare the full operating cost rather than only the license. Measure response latency, context length, concurrent users, tool-call reliability, structured-output compliance and task-specific accuracy on the hardware and runtime you actually plan to use.

Bottom line

gpt-oss-20b is a downloadable OpenAI reasoning model designed to make local and private deployment practical at a relatively modest hardware scale. Its 128K context, configurable reasoning effort, tool-use support, structured-output compatibility and fine-tuning path make it suitable for developers who want control over the model and serving environment.

Its central trade-off is clear: the model avoids an official per-token API bill and gives operators control, but it shifts infrastructure, security and reliability responsibilities to them. It is therefore best viewed as a self-hosted reasoning foundation for text-based applications, not as a direct replacement for a managed multimodal service.


Answers to Frequently Asked Questions

How much does gpt-oss-20b cost, and what license does it use?
There is no official OpenAI per-token API price for gpt-oss-20b because OpenAI does not host it through the OpenAI API. The weights are available under the Apache 2.0 license, subject to OpenAI's gpt-oss usage policy. Although the model is free to download, deployment can incur costs for hardware, cloud infrastructure, storage, electricity, networking, maintenance, and engineering.
Is gpt-oss-20b available through the OpenAI API or ChatGPT?
No. gpt-oss-20b is not included in ChatGPT and is not offered as an official model through the OpenAI API. Developers must download and operate the weights themselves or use a third-party hosted inference provider.
What hardware is required to run gpt-oss-20b locally?
OpenAI states that gpt-oss-20b can run within approximately 16 GB of memory using its MXFP4-quantized checkpoint. The actual requirement may be higher because of runtime overhead, operating-system memory, long contexts, key-value cache usage, and concurrent requests. Efficient serving also depends on the selected hardware and inference runtime.
What is gpt-oss-20b?
gpt-oss-20b is an open-weight reasoning language model released by OpenAI on August 5, 2025. It has approximately 21 billion total parameters, uses a mixture-of-experts architecture with about 3.6 billion active parameters per token, and can be downloaded and run locally, on-premises, in a private cloud, or through third-party inference providers.
How much context does gpt-oss-20b support?
gpt-oss-20b supports a documented maximum context length of 128K tokens, or 131,072 tokens. This context includes instructions, conversation history, documents, tool definitions, and the generated response. Actual performance and memory usage depend on the runtime, hardware, context size, and number of concurrent requests.


Sources 7
Provider

About OpenAI