GPT-OSS

gpt-oss-120b

by OpenAI · Current open-weight model; downloadable and deployable locally or through third-party providers; not available through the OpenAI API

OpenAI's GPT-OSS-120B is a downloadable, Apache 2.0-licensed reasoning model with approximately 117 billion total parameters, 5.1 billion active parameters per token, a 131,072-token context window, configurable reasoning effort, tool use, structured outputs, and fine-tuning support. It is text-only and intended primarily for self-hosted or third-party deployment.

Text Reasoning Coding
GPT-OSS-120B is OpenAI's most capable open-weight language model. Released on August 5, 2025, it is designed for reasoning, coding, agentic workflows, and customizable deployment rather than direct hosted use through the OpenAI API. Its mixture-of-experts architecture activates approximately 5.1 billion parameters per token and can fit into an 80GB-class GPU under suitable deployment conditions.
Outputs

What gpt-oss-120b can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Fine-tuning Structured output
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
6/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family GPT-OSS
Model type Reasoning
Context window 131K tokens
Maximum output 131K tokens
Knowledge cutoff 2024-06-01
Release date 2025-08-05
Status Current open-weight model; downloadable and deployable locally or through third-party providers; not available through the OpenAI API
Knowledge cutoff notes

OpenAI's current developer model page specifies June 1, 2024 as the knowledge cutoff. Web search or retrieval tools can provide newer information during use but do not change the underlying cutoff.

Model notes

OpenAI released gpt-oss-120b on August 5, 2025 under the Apache 2.0 license. It has approximately 117 billion total parameters and activates about 5.1 billion parameters per token through a mixture-of-experts architecture. The model supports low, medium, and high reasoning effort, function calling, structured outputs, and tool-oriented workflows such as web search and Python execution when connected to compatible tooling. It is text-only and does not natively accept or generate image, audio, or video content. OpenAI's help documentation states that the model is not served through the OpenAI API, so there is no direct OpenAI API token price; deployment costs depend on hardware or third-party inference providers. OpenAI's developer model page lists a 131,072-token context window, a 131,072-token maximum output, June 1, 2024 knowledge cutoff, streaming, function calling, structured outputs, and multiple API endpoint compatibility entries, but the listed rate limits are zero. The model is distributed in native MXFP4 quantization and is designed to fit within an 80GB-class GPU under suitable deployment conditions. Editorial scores are comparative estimates, not official vendor ratings.

Model guide

GPT-OSS-120B: Features, Context, License and Deployment

GPT-OSS-120B is OpenAI's largest open-weight reasoning model, with approximately 117 billion total parameters, a mixture-of-experts architecture, a 131,072-token context window, configurable reasoning effort, tool use, structured outputs, and Apache 2.0 licensing for local or third-party deployment.

What is GPT-OSS-120B?

GPT-OSS-120B is OpenAI's largest open-weight reasoning model. OpenAI released it on August 5, 2025, alongside the smaller GPT-OSS-20B. Unlike models accessed only through a provider-managed endpoint, GPT-OSS-120B can be downloaded, customized, fine-tuned, and deployed on infrastructure operated by a developer, organization, or third-party inference provider.

The model is aimed at applications that need reasoning, coding, instruction following, and multi-step agent workflows. It is distributed under the Apache 2.0 license, which generally permits commercial use, modification, and redistribution subject to the license terms. OpenAI positions the model as an option for teams that want more control over deployment, data handling, and model customization than a conventional hosted API provides.

GPT-OSS-120B is text-only. It produces text or structured textual responses, although an application can connect those responses to external tools such as web search or Python execution. The model itself does not natively generate images, audio, video, or other non-text media.

Architecture, parameters, and context limits

GPT-OSS-120B uses a mixture-of-experts, or MoE, Transformer architecture. It contains approximately 117 billion total parameters, but only about 5.1 billion are activated for each token. In practical terms, the model has a large overall capacity while using a smaller active computation path for each part of a response.

The architecture has 36 layers and 128 experts, with four experts activated per token. This sparse design can reduce inference computation compared with a dense model containing a similar number of total parameters, although actual speed and memory use depend on the serving runtime, quantization, hardware, batch size, and context length.

OpenAI's current model documentation lists a context window of 131,072 tokens and a maximum output length of up to 131,072 tokens. The context window includes the material supplied to the model and the generated response, so very long prompts can reduce the space available for output. Real-world deployments may impose lower limits because of available memory, service configuration, or provider-specific restrictions.

OpenAI states that the model can fit within a single 80GB-class GPU, such as an H100, when using its native MXFP4 quantization. This is a deployment target rather than a guarantee for every configuration. Longer sequences, concurrent requests, different quantization choices, and production serving overhead can require additional memory or more than one GPU.

Reasoning and agentic capabilities

GPT-OSS-120B supports low, medium, and high reasoning effort settings. Lower effort can reduce latency and computational cost, while higher effort gives the model more opportunity to work through difficult problems. The most suitable setting depends on whether the application values response time, infrastructure cost, or deeper problem solving.

The model was post-trained using methods associated with reasoning-model development, including supervised fine-tuning and reinforcement learning. OpenAI presents it for tasks involving multi-step reasoning, coding, scientific and mathematical work, instruction following, and tool-oriented agents.

Its agentic capabilities do not mean that the model independently performs every external action. The model can produce a function call or structured request, but the surrounding application must validate that request, execute the selected tool, and return the result. This distinction matters for security: developers remain responsible for authorization, input validation, tool permissions, logging, and defenses against prompt injection.

Coding and evaluation profile

OpenAI says GPT-OSS-120B was trained with an emphasis on STEM, coding, and general knowledge. The company reports strong results in reasoning, competition mathematics, coding, and tool-calling evaluations. OpenAI also states that the model outperformed o3-mini and matched or exceeded o4-mini on selected coding, general problem-solving, and tool-use benchmarks, with particularly strong results on competition mathematics and health-related evaluations.

Those are provider-reported benchmark claims, not guarantees for an individual application. Results can change with the reasoning setting, prompt format, chat template, runtime, quantization, tool implementation, and any fine-tuning performed by the deployer. A team evaluating the model should test representative tasks, including failure cases, latency, memory consumption, and the accuracy of tool calls.

For coding assistants, GPT-OSS-120B is most useful when it can inspect relevant code, reason over several files, and use controlled tools for tests or documentation lookup. It should not be given unrestricted ability to modify production systems without review and environment-level safeguards.

How GPT-OSS-120B is deployed

The model weights are available through OpenAI's Hugging Face repository. OpenAI also provides ecosystem guidance for tools such as Transformers, vLLM, Ollama, and LM Studio. The model was trained for OpenAI's Harmony response format, so an implementation should use a compatible chat template or renderer rather than assuming that a generic chat prompt will provide optimal behavior.

GPT-OSS-120B is primarily a self-hosted or third-party-hosted model. OpenAI's help documentation states that it is not available through the OpenAI API. As a result, there is no official OpenAI-hosted per-token price for this model. The cost of using it depends on the selected hardware, cloud GPU rental or inference provider, quantization, traffic volume, storage, orchestration, and operational requirements.

This pricing model is fundamentally different from a standard hosted API. Self-hosting can provide greater control over data and customization, but the operator must pay for infrastructure and manage scaling, uptime, updates, monitoring, and security. A third-party provider may simplify operations while introducing provider-specific prices, limits, privacy terms, and availability.

The Apache 2.0 license is an important advantage for organizations that need commercial-friendly redistribution or modification rights. However, licensing does not remove the need for responsible deployment. Operators remain accountable for privacy, access control, content policies, monitoring, and the risks created by applications built around the model.

Modalities, structured output, and knowledge cutoff

GPT-OSS-120B accepts text input and returns text output or structured textual output. It does not natively accept images, audio, or video, and it does not generate images, audio, music, or video. An application that needs vision, speech, or media generation must combine GPT-OSS-120B with separate models or processing components.

The model supports structured outputs and function calling when integrated with compatible tooling. Structured output can make responses easier for software to parse, but it should not be confused with the model independently carrying out an action. Applications should still validate generated fields and reject malformed, unsafe, or unauthorized requests.

OpenAI lists June 1, 2024 as the knowledge cutoff. Web search or retrieval tools can provide newer information during a session, but retrieval does not change the model's underlying training cutoff. For current events, changing documentation, prices, or live operational data, an application should use an appropriate retrieval or search component and clearly distinguish retrieved information from the model's prior knowledge.

Main strengths and limitations

  • Open deployment: Downloadable weights allow local, private, customized, or third-party deployment instead of requiring OpenAI-hosted inference.
  • Reasoning control: Low, medium, and high reasoning effort settings provide a practical latency and quality trade-off.
  • Large context: The documented 131,072-token context window supports long documents, codebases, and multi-step conversations when hardware permits.
  • Tool-oriented design: Function calling, structured outputs, and agent workflows support applications that connect the model to controlled external tools.
  • Commercial-friendly license: Apache 2.0 licensing supports modification and commercial use subject to its terms.
  • Infrastructure responsibility: Self-hosting requires GPU capacity, serving expertise, monitoring, security controls, and maintenance.
  • No native media support: The model is not suitable by itself for image, audio, video, or multimodal input and output.
  • Knowledge cutoff: Without retrieval, the model cannot be expected to know events or changes after June 1, 2024.
  • Limited central control after download: Once deployed, the weights cannot receive automatic provider-side safety or behavior updates in the same way as a centrally hosted service.

When to choose GPT-OSS-120B

Choose GPT-OSS-120B when deployment control is more important than turnkey hosting. It is a reasonable candidate for private enterprise assistants, coding systems, research projects, custom domain adaptation, evaluation environments, and agentic applications that need local or third-party inference.

It is especially attractive when an organization wants to inspect or modify the model, fine-tune it, control where data is processed, or avoid depending on a single provider's hosted API. Its large context window and configurable reasoning effort can also help with long documents, complex coding tasks, and multi-step workflows.

A hosted API model may be more appropriate when the priority is quick integration, predictable provider-managed operations, automatic infrastructure scaling, or access to native vision, audio, and other modalities. A smaller model may be preferable when low latency, low hardware cost, or high request volume matters more than maximum reasoning capacity. Conversely, a specialized multimodal or audio model is a better fit when text-only processing is insufficient.

GPT-OSS-120B should therefore be evaluated as an open-weight deployment choice, not simply as a lower-priced version of an OpenAI-hosted model. Its key trade-off is control and customizability versus the engineering work and infrastructure cost required to operate it.

Bottom line

GPT-OSS-120B combines a large sparse mixture-of-experts architecture with a 131,072-token context window, configurable reasoning effort, coding and tool-use support, structured outputs, and Apache 2.0 licensing. Its strongest use cases are self-hosted or third-party deployments where teams need control, customization, and agentic reasoning.

The model is less suitable for users seeking a managed OpenAI API endpoint, built-in multimodal capabilities, minimal infrastructure work, or provider-operated safety updates. Its actual value depends not only on model quality but also on GPU availability, serving configuration, tool design, evaluation discipline, and the operational capabilities of the deploying organization.


Answers to Frequently Asked Questions

Does GPT-OSS-120B support images, audio, or video?
No. GPT-OSS-120B is a text-only model that accepts text and produces text or structured textual output. Applications requiring vision, speech, image generation, audio, or video must combine it with separate models or processing components.
What license does GPT-OSS-120B use?
GPT-OSS-120B is distributed under the Apache 2.0 license, which generally permits commercial use, modification, and redistribution subject to the license terms. Organizations must still manage privacy, security, access control, monitoring, and other deployment responsibilities.
Is GPT-OSS-120B available through the OpenAI API?
No. GPT-OSS-120B is not available through the OpenAI API. It is intended for self-hosted or third-party-hosted deployment, so users must account for GPU infrastructure, inference provider fees, storage, scaling, monitoring, and security.
What are GPT-OSS-120B's context window and hardware requirements?
GPT-OSS-120B has a documented context window of 131,072 tokens and a maximum output length of up to 131,072 tokens. OpenAI states that it can fit on a single 80GB-class GPU, such as an H100, when using native MXFP4 quantization, although longer sequences, concurrent requests, and production overhead may require additional memory or GPUs.
What is GPT-OSS-120B?
GPT-OSS-120B is OpenAI's largest open-weight reasoning model, released on August 5, 2025. It is designed for reasoning, coding, instruction following, and multi-step agent workflows, and can be downloaded, customized, fine-tuned, and deployed on self-managed or third-party infrastructure.


Sources 6
Provider

About OpenAI