What is NVIDIA Nemotron 3 Super 120B-A12B?
NVIDIA Nemotron 3 Super 120B-A12B is an open-weight language model from NVIDIA's Nemotron 3 family. Its name describes a model with 120 billion total parameters and approximately 12 billion active parameters during inference. It uses a mixture-of-experts design, meaning that different inputs can be routed through selected portions of the network rather than activating every parameter for every token.
For practical users, this makes the model a better fit for demanding text workloads than for everyday lightweight chat. Its stated focus includes multi-step reasoning, software development, tool-using agents, retrieval-augmented generation (RAG), long documents, and high-volume enterprise inference. The model produces text only: it does not accept images, audio, or video and does not generate those media types.
NVIDIA provides the model as downloadable checkpoints as well as through NVIDIA NIM and a hosted API trial endpoint. This gives organizations a choice between using an NVIDIA-managed interface and deploying the weights in infrastructure they control. The available checkpoint formats include BF16, FP8, and NVFP4, with both post-trained and base BF16 checkpoints documented by NVIDIA.
Architecture and 1-million-token context
Nemotron 3 Super combines Mamba-2, latent mixture-of-experts, and attention components in a hybrid architecture. It also uses multi-token prediction. These are implementation details rather than features a user selects in a chat window, but they help explain the model's positioning: it is designed to handle long sequences and substantial inference workloads while using a relatively small active parameter count compared with its total size.
The documented context length is up to 1,000,000 tokens. A context window is the amount of input and generated conversation history that the model can process in one request, although the exact usable amount can depend on the serving configuration and how many output tokens are reserved. A million-token context can support very large document collections, long code repositories, extended agent traces, or substantial RAG inputs.
NVIDIA's documentation does not clearly specify a single maximum generated-output-token limit for this model. The 1-million-token figure should therefore not be interpreted as a guaranteed output length. It describes the supported context capacity, while the practical output limit may depend on the NIM, vLLM, hosted endpoint, or other deployment configuration.
Reasoning, coding, and tool use
Reasoning can be enabled or disabled through the model's chat template. This is useful when an application needs a deliberate multi-step response for a difficult task, but wants lower latency or shorter answers for routine requests. NVIDIA recommends a temperature of 1.0 and top_p of 0.95 for reasoning, tool calling, and general chat according to the supplied model documentation.
In editorial evaluation, the model is best characterized as having strong reasoning and good coding capability, with an assessment of 9 out of 10 for reasoning and 8 out of 10 for coding. These are editorial scores, not NVIDIA-published benchmark results. They indicate the intended practical positioning rather than a standardized performance guarantee.
The model supports tool calling with JSON arguments through compatible NIM and vLLM deployments. An agent can therefore ask it to select a function, provide arguments, receive the tool result, and continue the conversation. This is relevant to workflows such as database queries, business-system actions, retrieval pipelines, calculators, and software automation. Tool support depends on the serving stack and application integration; it does not mean the model independently has access to the web, private databases, or external software.
Nemotron 3 Super does not provide built-in web search. It also does not have a documented native image, audio, or video input mode. A text-based RAG system can supply extracted document content, but that is different from direct multimodal understanding.
Where it fits in NVIDIA's model catalog
Within NVIDIA's current AI ecosystem, Nemotron 3 Super is a model for developers and organizations that need control over deployment, inference infrastructure, and agent behavior. It is not a consumer chatbot subscription or a small local assistant intended to run on ordinary hardware.
NVIDIA makes the model available through NIM, its inference-microservice approach, and through downloadable weights for self-hosted use. The model is also listed in NVIDIA's model and API catalogs, with documentation for NIM and NVIDIA Dynamo deployment. This positioning makes it particularly relevant to data-center, workstation, and enterprise environments where teams can provision substantial GPU capacity and tune serving performance.
The open-weight availability is important, but it should not be confused with zero-cost operation. Downloading a checkpoint may avoid a per-token provider charge, yet self-hosting still requires suitable GPUs, storage, networking, maintenance, and engineering work. NVIDIA's supplied information identifies the weights as available under the NVIDIA Nemotron Open Model License.
Main strengths and trade-offs
- Very large context: The supported context length of up to 1 million tokens is useful for long documents, codebases, RAG collections, and extended agent sessions.
- Reasoning control: Applications can enable or disable reasoning through the chat template instead of treating every request as an equally expensive reasoning task.
- Agent support: Compatible deployments can use tool calling with JSON arguments, making the model suitable for function-driven workflows.
- Deployment flexibility: NVIDIA offers hosted access, NIM deployment, and downloadable BF16, FP8, and NVFP4 checkpoints.
- Efficient active computation: Although the model has 120 billion total parameters, approximately 12 billion are active for a given inference path, reflecting its mixture-of-experts design.
The main trade-off is infrastructure complexity. A 120-billion-parameter model is not a lightweight option, even when only a fraction of parameters is active for each token. The required hardware and serving configuration can make it unsuitable for small teams, modest workstations, or applications where a smaller model provides adequate quality.
Speed and cost also depend heavily on deployment. The supplied assessment rates speed at 8 out of 10 and cost efficiency at 7 out of 10, but these are editorial evaluations rather than fixed provider specifications. Quantization formats such as FP8 and NVFP4 may help reduce the resource burden, while BF16 can require more memory. Actual throughput, latency, and total cost vary with GPU type, batching, context size, concurrency, and software configuration.
Pricing and access
No official per-token hosted API price was identified in the supplied documentation. The model is described as available through NVIDIA's hosted API trial endpoint, NIM, downloadable checkpoints, and self-hosted deployments, but a generally applicable recurring or per-token price cannot be stated reliably.
For a self-hosted installation, the relevant cost is infrastructure rather than a published model subscription. Organizations should account for GPU acquisition or rental, storage for the checkpoints, deployment engineering, monitoring, scaling, and ongoing operations. Hosted access may be simpler, but users should confirm the current endpoint's quotas, availability, and pricing before building a production workflow.
Best use cases
Nemotron 3 Super is a strong candidate when an application needs several of the following characteristics:
- Long-context analysis across large document sets or repositories.
- Agentic workflows in which the model plans steps and calls external tools.
- Complex reasoning with an option to turn reasoning off for simpler requests.
- Code generation, code explanation, debugging, and software-engineering assistance.
- RAG systems that need to pass substantial retrieved context to the model.
- Private or controlled deployments using downloadable weights and NVIDIA infrastructure.
- High-volume inference where an organization can optimize batching and serving.
For example, an enterprise could use the model to examine a large collection of technical documents, retrieve relevant passages, call an internal search or ticketing function, and produce a grounded response. A development platform could use it to inspect a large codebase and invoke testing or repository tools. These examples require application-level integrations; the model does not automatically connect to those systems.
When another option may be more appropriate
A smaller language model may be a better choice for short chats, simple extraction, low-latency classification, or deployment on limited hardware. Nemotron 3 Super's total model size and long-context capability can be unnecessary overhead when requests are short and straightforward.
A multimodal model is more appropriate when the application must directly interpret images, audio, or video. Nemotron 3 Super is text-only for both input and output according to the supplied specifications. Similarly, a service with documented web search should be preferred when current web information is a core requirement, because this model does not provide native web search.
Teams that need a predictable public per-token price should also compare hosted alternatives with clearly published pricing. NVIDIA's supplied information does not establish a standard price for this model, so cost planning may be easier with a provider that publishes fixed input and output rates.
Overall assessment
NVIDIA Nemotron 3 Super 120B-A12B is aimed at serious long-context and agentic workloads rather than casual consumer use. Its combination of a 1-million-token context, configurable reasoning, tool calling, open-weight checkpoints, and multiple deployment paths makes it appealing to teams that need control and can support substantial infrastructure.
Its limitations are equally important: it is text-only, has no native web search, has no clearly documented maximum output-token figure, and has no verified general per-token price in the supplied research. The model makes the most sense when long context, reasoning, coding, and self-managed or NVIDIA-centered deployment justify the operational cost. For smaller, cheaper, faster, or multimodal applications, another type of model may be a better fit.

