Nemotron 3

NVIDIA Nemotron 3 Super 120B-A12B

by NVIDIA AI · Current; open-weight; available through NVIDIA NIM, NVIDIA's hosted API trial endpoint, downloadable checkpoints, and self-hosted deployments.

An open-weight NVIDIA language model with 120 billion total parameters, approximately 12 billion active parameters, hybrid Mamba-Transformer architecture, configurable reasoning, tool calling, and up to 1 million tokens of context.

Text Reasoning Coding
NVIDIA Nemotron 3 Super 120B-A12B is built for agentic reasoning, coding, long-context analysis, retrieval-augmented generation, conversational workloads, and high-volume inference. It is available through NVIDIA NIM and a hosted API trial endpoint, while NVIDIA also provides downloadable BF16, FP8, and NVFP4 checkpoints for compatible deployments.
Outputs

What NVIDIA Nemotron 3 Super 120B-A12B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning
Model profile

Performance characteristics

9/10 Reasoning
8/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Nemotron 3
Model type Reasoning
Context window 1M tokens
Knowledge cutoff February 2026 for post-training data; June 2025 for pre-training data
Release date 2026-03-11
Status Current; open-weight; available through NVIDIA NIM, NVIDIA's hosted API trial endpoint, downloadable checkpoints, and self-hosted deployments.
Knowledge cutoff notes

NVIDIA's NGC model card distinguishes the data freshness of the post-training and pre-training stages rather than providing one single cutoff date. Post-training data is reported through February 2026, while pre-training data is reported through June 2025.

Model notes

The model has 120B total parameters and approximately 12B active parameters. It uses a hybrid Mamba-2, latent mixture-of-experts, and attention architecture with multi-token prediction. NVIDIA publishes post-trained NVFP4, FP8, and BF16 checkpoints plus a base BF16 checkpoint. Reasoning can be enabled or disabled through the chat template, and NVIDIA recommends temperature 1.0 and top_p 0.95 for reasoning, tool calling, and general chat. Tool calling with JSON arguments is supported through compatible NIM and vLLM deployments. The model card reports post-training data through February 2026 and pre-training data through June 2025. Official per-token hosted API pricing was not identified; downloadable weights and self-hosted deployment are available under the NVIDIA Nemotron Open Model License. The exact maximum generated-output-token limit is not clearly specified in the authoritative model documentation reviewed.

Model guide

NVIDIA Nemotron 3 Super 120B-A12B: Open-Weight Reasoning for Long-Context Agents

NVIDIA Nemotron 3 Super 120B-A12B is an open-weight 120-billion-parameter mixture-of-experts language model with approximately 12 billion active parameters, a hybrid Mamba-Transformer architecture, native reasoning modes, tool calling, and support for context windows up to 1 million tokens.

What is NVIDIA Nemotron 3 Super 120B-A12B?

NVIDIA Nemotron 3 Super 120B-A12B is an open-weight language model from NVIDIA's Nemotron 3 family. Its name describes a model with 120 billion total parameters and approximately 12 billion active parameters during inference. It uses a mixture-of-experts design, meaning that different inputs can be routed through selected portions of the network rather than activating every parameter for every token.

For practical users, this makes the model a better fit for demanding text workloads than for everyday lightweight chat. Its stated focus includes multi-step reasoning, software development, tool-using agents, retrieval-augmented generation (RAG), long documents, and high-volume enterprise inference. The model produces text only: it does not accept images, audio, or video and does not generate those media types.

NVIDIA provides the model as downloadable checkpoints as well as through NVIDIA NIM and a hosted API trial endpoint. This gives organizations a choice between using an NVIDIA-managed interface and deploying the weights in infrastructure they control. The available checkpoint formats include BF16, FP8, and NVFP4, with both post-trained and base BF16 checkpoints documented by NVIDIA.

Architecture and 1-million-token context

Nemotron 3 Super combines Mamba-2, latent mixture-of-experts, and attention components in a hybrid architecture. It also uses multi-token prediction. These are implementation details rather than features a user selects in a chat window, but they help explain the model's positioning: it is designed to handle long sequences and substantial inference workloads while using a relatively small active parameter count compared with its total size.

The documented context length is up to 1,000,000 tokens. A context window is the amount of input and generated conversation history that the model can process in one request, although the exact usable amount can depend on the serving configuration and how many output tokens are reserved. A million-token context can support very large document collections, long code repositories, extended agent traces, or substantial RAG inputs.

NVIDIA's documentation does not clearly specify a single maximum generated-output-token limit for this model. The 1-million-token figure should therefore not be interpreted as a guaranteed output length. It describes the supported context capacity, while the practical output limit may depend on the NIM, vLLM, hosted endpoint, or other deployment configuration.

Reasoning, coding, and tool use

Reasoning can be enabled or disabled through the model's chat template. This is useful when an application needs a deliberate multi-step response for a difficult task, but wants lower latency or shorter answers for routine requests. NVIDIA recommends a temperature of 1.0 and top_p of 0.95 for reasoning, tool calling, and general chat according to the supplied model documentation.

In editorial evaluation, the model is best characterized as having strong reasoning and good coding capability, with an assessment of 9 out of 10 for reasoning and 8 out of 10 for coding. These are editorial scores, not NVIDIA-published benchmark results. They indicate the intended practical positioning rather than a standardized performance guarantee.

The model supports tool calling with JSON arguments through compatible NIM and vLLM deployments. An agent can therefore ask it to select a function, provide arguments, receive the tool result, and continue the conversation. This is relevant to workflows such as database queries, business-system actions, retrieval pipelines, calculators, and software automation. Tool support depends on the serving stack and application integration; it does not mean the model independently has access to the web, private databases, or external software.

Nemotron 3 Super does not provide built-in web search. It also does not have a documented native image, audio, or video input mode. A text-based RAG system can supply extracted document content, but that is different from direct multimodal understanding.

Where it fits in NVIDIA's model catalog

Within NVIDIA's current AI ecosystem, Nemotron 3 Super is a model for developers and organizations that need control over deployment, inference infrastructure, and agent behavior. It is not a consumer chatbot subscription or a small local assistant intended to run on ordinary hardware.

NVIDIA makes the model available through NIM, its inference-microservice approach, and through downloadable weights for self-hosted use. The model is also listed in NVIDIA's model and API catalogs, with documentation for NIM and NVIDIA Dynamo deployment. This positioning makes it particularly relevant to data-center, workstation, and enterprise environments where teams can provision substantial GPU capacity and tune serving performance.

The open-weight availability is important, but it should not be confused with zero-cost operation. Downloading a checkpoint may avoid a per-token provider charge, yet self-hosting still requires suitable GPUs, storage, networking, maintenance, and engineering work. NVIDIA's supplied information identifies the weights as available under the NVIDIA Nemotron Open Model License.

Main strengths and trade-offs

  • Very large context: The supported context length of up to 1 million tokens is useful for long documents, codebases, RAG collections, and extended agent sessions.
  • Reasoning control: Applications can enable or disable reasoning through the chat template instead of treating every request as an equally expensive reasoning task.
  • Agent support: Compatible deployments can use tool calling with JSON arguments, making the model suitable for function-driven workflows.
  • Deployment flexibility: NVIDIA offers hosted access, NIM deployment, and downloadable BF16, FP8, and NVFP4 checkpoints.
  • Efficient active computation: Although the model has 120 billion total parameters, approximately 12 billion are active for a given inference path, reflecting its mixture-of-experts design.

The main trade-off is infrastructure complexity. A 120-billion-parameter model is not a lightweight option, even when only a fraction of parameters is active for each token. The required hardware and serving configuration can make it unsuitable for small teams, modest workstations, or applications where a smaller model provides adequate quality.

Speed and cost also depend heavily on deployment. The supplied assessment rates speed at 8 out of 10 and cost efficiency at 7 out of 10, but these are editorial evaluations rather than fixed provider specifications. Quantization formats such as FP8 and NVFP4 may help reduce the resource burden, while BF16 can require more memory. Actual throughput, latency, and total cost vary with GPU type, batching, context size, concurrency, and software configuration.

Pricing and access

No official per-token hosted API price was identified in the supplied documentation. The model is described as available through NVIDIA's hosted API trial endpoint, NIM, downloadable checkpoints, and self-hosted deployments, but a generally applicable recurring or per-token price cannot be stated reliably.

For a self-hosted installation, the relevant cost is infrastructure rather than a published model subscription. Organizations should account for GPU acquisition or rental, storage for the checkpoints, deployment engineering, monitoring, scaling, and ongoing operations. Hosted access may be simpler, but users should confirm the current endpoint's quotas, availability, and pricing before building a production workflow.

Best use cases

Nemotron 3 Super is a strong candidate when an application needs several of the following characteristics:

  • Long-context analysis across large document sets or repositories.
  • Agentic workflows in which the model plans steps and calls external tools.
  • Complex reasoning with an option to turn reasoning off for simpler requests.
  • Code generation, code explanation, debugging, and software-engineering assistance.
  • RAG systems that need to pass substantial retrieved context to the model.
  • Private or controlled deployments using downloadable weights and NVIDIA infrastructure.
  • High-volume inference where an organization can optimize batching and serving.

For example, an enterprise could use the model to examine a large collection of technical documents, retrieve relevant passages, call an internal search or ticketing function, and produce a grounded response. A development platform could use it to inspect a large codebase and invoke testing or repository tools. These examples require application-level integrations; the model does not automatically connect to those systems.

When another option may be more appropriate

A smaller language model may be a better choice for short chats, simple extraction, low-latency classification, or deployment on limited hardware. Nemotron 3 Super's total model size and long-context capability can be unnecessary overhead when requests are short and straightforward.

A multimodal model is more appropriate when the application must directly interpret images, audio, or video. Nemotron 3 Super is text-only for both input and output according to the supplied specifications. Similarly, a service with documented web search should be preferred when current web information is a core requirement, because this model does not provide native web search.

Teams that need a predictable public per-token price should also compare hosted alternatives with clearly published pricing. NVIDIA's supplied information does not establish a standard price for this model, so cost planning may be easier with a provider that publishes fixed input and output rates.

Overall assessment

NVIDIA Nemotron 3 Super 120B-A12B is aimed at serious long-context and agentic workloads rather than casual consumer use. Its combination of a 1-million-token context, configurable reasoning, tool calling, open-weight checkpoints, and multiple deployment paths makes it appealing to teams that need control and can support substantial infrastructure.

Its limitations are equally important: it is text-only, has no native web search, has no clearly documented maximum output-token figure, and has no verified general per-token price in the supplied research. The model makes the most sense when long context, reasoning, coding, and self-managed or NVIDIA-centered deployment justify the operational cost. For smaller, cheaper, faster, or multimodal applications, another type of model may be a better fit.


Answers to Frequently Asked Questions

Is NVIDIA Nemotron 3 Super 120B-A12B multimodal, and does it include web search?
No. NVIDIA Nemotron 3 Super 120B-A12B is a text-only model for both input and output, with no documented native image, audio, or video support. It also does not provide built-in web search, although an application can connect it to external search or RAG systems.
How can NVIDIA Nemotron 3 Super 120B-A12B be deployed?
NVIDIA provides hosted API access, NIM deployment, and downloadable checkpoints for self-hosting. Available checkpoint formats include BF16, FP8, and NVFP4. Self-hosting requires substantial GPU capacity, storage, networking, deployment engineering, and ongoing operations.
Does NVIDIA Nemotron 3 Super 120B-A12B support reasoning and tool calling?
Yes. Reasoning can be enabled or disabled through the model's chat template, and compatible NIM and vLLM deployments support tool calling with JSON arguments. The model can therefore participate in workflows involving databases, retrieval systems, business applications, calculators, and software tools.
How large is the context window of NVIDIA Nemotron 3 Super 120B-A12B?
The model supports a context length of up to 1,000,000 tokens. This can accommodate large document collections, code repositories, RAG inputs, and extended agent traces, although the usable amount depends on the serving configuration and output-token allocation.
What is NVIDIA Nemotron 3 Super 120B-A12B?
NVIDIA Nemotron 3 Super 120B-A12B is an open-weight mixture-of-experts language model with 120 billion total parameters and approximately 12 billion active parameters during inference. It is designed for reasoning, coding, long-context analysis, retrieval-augmented generation, and tool-using agents.


Sources 6
Provider

About NVIDIA AI