What is gpt-oss-20b?
gpt-oss-20b is an open-weight reasoning language model provided by OpenAI and released on August 5, 2025. “Open-weight” means that the model weights can be downloaded and run by developers instead of being available only behind a managed provider endpoint. Users can deploy it locally, on-premises, in a private cloud or through a third-party inference host.
The model has approximately 21 billion total parameters. It uses a mixture-of-experts architecture, so only a subset of those parameters is active for each token. OpenAI documents approximately 3.6 billion active parameters per token. This design aims to provide a larger model's representational capacity while reducing the computation required for each generated token.
gpt-oss-20b is the smaller model in the original gpt-oss release, alongside gpt-oss-120b. The two models share the open-weight deployment approach, but gpt-oss-20b is positioned for lower-latency and more hardware-conscious deployments. It is not included in ChatGPT and is not served as an official model through the OpenAI API.
Who should use gpt-oss-20b?
The model is primarily intended for developers and organizations that need more control over inference than a hosted API normally provides. Running the weights yourself can support private processing, custom safety policies, specialized tool integrations and fine-tuning with organization-specific data.
It is especially relevant for coding assistants, internal question-answering systems, local research tools, agentic workflows and applications that need a reasoning model without sending every prompt to a provider-managed service. It can also be useful for rapid experimentation because the weights and reference integrations are available through the OpenAI model repository and related open-source tooling.
This flexibility comes with operational responsibility. A self-hosted deployment requires decisions about hardware, serving software, access control, monitoring, logging, abuse prevention, updates and evaluation. The model is therefore not a drop-in replacement for a managed API if the main priority is immediate production deployment with minimal infrastructure work.
Architecture and 128K context window
gpt-oss-20b uses a 24-layer Transformer architecture with mixture-of-experts routing. The configuration includes 32 local experts and activates four experts per token. It also uses alternating sliding and full attention patterns, grouped multi-query attention and rotary positional embeddings. These details affect how the model processes long inputs and how efficiently an inference runtime can serve it.
The documented maximum context length is 128K tokens, or 131,072 tokens in the model configuration. The context window includes the information supplied to the model during a request, such as system instructions, conversation history, documents, tool definitions and the generated response. A long context limit does not guarantee that every runtime will perform equally well at the maximum length; memory use and latency depend on context size, batching, hardware and serving software.
The released checkpoint uses MXFP4 quantization for the mixture-of-experts weights. Quantization stores model values in a more compact numerical format, reducing memory requirements compared with higher-precision weights. OpenAI states that gpt-oss-20b can run within approximately 16 GB of memory, but this is a target rather than a universal hardware requirement. Runtime overhead, operating-system memory, context length, concurrent requests and implementation details can increase the actual requirement.
Reasoning, coding and general capabilities
gpt-oss-20b is designed as a reasoning model rather than a basic text-completion model. It supports low, medium and high reasoning-effort settings. These settings let an application trade deliberation depth against response speed: lower effort can reduce latency, while higher effort may be more suitable for difficult multi-step tasks. The exact quality and speed difference depends on the prompt, runtime and hardware.
The model was trained for instruction following, STEM tasks, coding, general knowledge and agentic tool-use scenarios. In practical terms, it can be used to explain technical material, generate or review code, analyze text, follow structured instructions and break a problem into multiple steps before producing an answer.
OpenAI reports that gpt-oss-20b performs similarly to OpenAI o3-mini on several published evaluations. OpenAI's comparison information lists scores of 85.3 on MMLU, 71.5 on GPQA Diamond, 17.3 on Humanity's Last Exam, 96.0 on AIME 2024 and 98.7 on AIME 2025. These are provider-reported benchmark results under the stated evaluation conditions. They are useful for understanding the intended positioning, but they should not be treated as a guarantee of performance on a particular application or dataset.
The model's coding capability is best understood as part of its broader reasoning and instruction-following design. It can support coding assistants, code explanation, debugging and tool-connected development workflows. The supplied research does not establish a universal programming-language ranking or a guaranteed software-engineering success rate, so production use should include repository-specific tests and human review.
Tools, function calling and structured output
gpt-oss-20b was trained for tool use and can be connected to function-calling workflows. A tool-connected application can ask the model to select an available function, provide arguments and incorporate the returned result into a later response. The model itself does not automatically gain access to the internet, databases or operating-system actions; those capabilities come from the tools and runtime that the operator connects.
OpenAI's reference ecosystem includes browser-based web browsing and Python execution examples. The model is also compatible with OpenAI-compatible local servers and several inference runtimes. Tool behavior, schemas, error handling and security controls can vary between Transformers, vLLM, SGLang, Ollama, LM Studio and other serving stacks.
Web search is an external capability rather than a change to the model's underlying knowledge. If a browser or search tool is connected, it supplies current information during inference. It does not establish a new knowledge cutoff or update the model weights.
The model uses OpenAI's Harmony response format. OpenAI recommends using this format for correct behavior when reasoning messages, tool calls or structured responses are involved. gpt-oss-20b also supports Structured Outputs when the serving implementation exposes that capability. Developers should verify the exact schema-enforcement behavior of their selected runtime instead of assuming that every OpenAI-compatible server provides identical guarantees.
Input and output modalities
gpt-oss-20b is text-only. It accepts text prompts and produces text responses, reasoning-related content, tool calls and structured machine-readable responses depending on the prompt format and serving implementation.
It does not natively accept images, audio or video. It also does not generate images, audio, music or video. An application can place extracted text or externally generated descriptions into a prompt, but that is not the same as native multimodal understanding. For image analysis, speech interaction or video generation, a model designed for those modalities is more appropriate.
Deployment options and hardware
OpenAI's official repository provides reference implementations and integrations for Transformers, vLLM, SGLang, Ollama, LM Studio, PyTorch, Triton and Apple Metal. The ecosystem also includes examples for downloading weights from Hugging Face and running the model through OpenAI-compatible local servers.
The approximately 16 GB memory target makes gpt-oss-20b more practical for some consumer GPUs, workstations and selected edge systems than larger open-weight models. However, “can run” and “can serve efficiently” are different requirements. A short single-user prompt may fit within a hardware target while long contexts, multiple simultaneous users, higher throughput or a large key-value cache require substantially more memory.
Runtime selection also affects speed. A well-optimized implementation with supported quantization and suitable hardware may deliver responsive local inference, while unsupported formats or constrained memory can cause slower generation, CPU offloading or reduced concurrency. Benchmarking the intended workload is more reliable than relying only on the nominal parameter count.
Pricing, license and API availability
There is no official OpenAI per-token input or output price for gpt-oss-20b because OpenAI does not host this model through the OpenAI API. The weights are available for download under the Apache 2.0 license, subject to OpenAI's gpt-oss usage policy.
“Free to download” does not mean that deployment has no cost. Operators may pay for a GPU or workstation, cloud instances, storage, electricity, networking, maintenance and engineering time. Third-party providers may offer hosted inference, but their prices depend on the provider, hardware, request volume, performance target and deployment configuration. Those prices should not be confused with an official OpenAI API rate.
The Apache 2.0 license generally supports commercial and private use subject to its terms, while the separate usage policy and applicable laws still matter. Organizations should review both before deploying the model in a product or regulated workflow.
Main strengths and limitations
Strengths
- Local control: The weights can be run on infrastructure managed by the developer or organization.
- Reasoning controls: Low, medium and high reasoning effort provide a practical latency-versus-deliberation trade-off.
- Long context: The documented 128K-token context supports lengthy instructions, documents and tool histories, subject to runtime memory limits.
- Efficient architecture: Mixture-of-experts routing activates approximately 3.6 billion parameters per token rather than all approximately 21 billion parameters.
- Customization: Developers can adjust prompts, tools, safety controls, runtimes and fine-tuning data.
- Open deployment ecosystem: Official references cover several widely used local and self-hosted runtimes.
Limitations
- No managed OpenAI endpoint: It is not available through the OpenAI API or ChatGPT.
- Text-only operation: Native image, audio and video input and output are not supported.
- Operational burden: The operator is responsible for scaling, security, monitoring, tool execution and safety evaluation.
- Variable runtime behavior: Streaming, tool use and structured outputs depend partly on the selected serving implementation.
- No documented maximum output limit in the supplied research: A provider-published maximum output-token value was not identified.
- Potentially variable quality: Benchmark results do not guarantee performance on a specific domain, codebase or business process.
When to choose gpt-oss-20b
Choose gpt-oss-20b when local or private inference is more important than access to a fully managed API. It is a sensible candidate for a coding assistant running near the development environment, an internal document assistant, a private agent with organization-controlled tools or an edge deployment where a larger model would be too slow or expensive.
It can also be attractive when an application needs to experiment with reasoning effort, function calling, structured responses or fine-tuning without committing to a provider-hosted model. Its smaller active parameter count and memory target may offer a better speed-and-cost profile than larger open-weight models, especially for single-user or moderate-throughput deployments.
Another option may be more appropriate when the application needs native image, audio or video capabilities, guaranteed managed availability, a documented hosted price, or minimal infrastructure ownership. A larger reasoning model may be preferable for tasks where maximum quality is more important than local resource requirements. Conversely, a smaller non-reasoning model may be preferable for simple classification or short responses where deliberation adds unnecessary latency.
For a deployment decision, compare the full operating cost rather than only the license. Measure response latency, context length, concurrent users, tool-call reliability, structured-output compliance and task-specific accuracy on the hardware and runtime you actually plan to use.
Bottom line
gpt-oss-20b is a downloadable OpenAI reasoning model designed to make local and private deployment practical at a relatively modest hardware scale. Its 128K context, configurable reasoning effort, tool-use support, structured-output compatibility and fine-tuning path make it suitable for developers who want control over the model and serving environment.
Its central trade-off is clear: the model avoids an official per-token API bill and gives operators control, but it shifts infrastructure, security and reliability responsibilities to them. It is therefore best viewed as a self-hosted reasoning foundation for text-based applications, not as a direct replacement for a managed multimodal service.

