Nemotron 3

NVIDIA Nemotron 3 Ultra 550B-A55B

by NVIDIA AI · Current; open-weight model with BF16 and NVFP4 checkpoints

NVIDIA Nemotron 3 Ultra 550B-A55B is a 550B-total, 55B-active open-weight text reasoning model for complex coding, long-context analysis, enterprise RAG, and tool-using agents. Its hybrid Mamba-2 and LatentMoE architecture supports configurable reasoning and up to 1M-token context through qualified serving profiles, but the model requires substantial multi-GPU infrastructure and has no verified provider-hosted per-token price.

Text Reasoning Coding
NVIDIA Nemotron 3 Ultra 550B-A55B is the largest model in NVIDIA's Nemotron 3 family. Released on June 4, 2026, it is an open-weight, text-only model intended for demanding reasoning, software engineering, research synthesis, long-context analysis, enterprise retrieval-augmented generation, and agentic workflows. The model combines a very large total parameter count with a sparse mixture-of-experts design, so only a portion of its parameters is active for each token. NVIDIA provides BF16 and NVFP4 checkpoints under the OpenMDW-1.1 license. Its headline context capability reaches up to 1 million tokens in supported serving configurations, although NVIDIA's deployment documentation distinguishes that profile from the model's native 256K configuration.
Outputs

What NVIDIA Nemotron 3 Ultra 550B-A55B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning
Model profile

Performance characteristics

10/10 Reasoning
9/10 Coding
7/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Nemotron 3
Model type Reasoning
Context window 1M tokens
Knowledge cutoff May 2026 post-training; September 2025 pre-training
Release date 2026-06-04
Status Current; open-weight model with BF16 and NVFP4 checkpoints
Knowledge cutoff notes

NVIDIA's model card states that post-training data has a cutoff date of May 2026 and pre-training data has a cutoff date of September 2025. These are separate data-freshness dates rather than a single unified knowledge-cutoff value.

Model notes

The exact model has 550B total parameters and up to 55B active parameters per token. NVIDIA documents up to 1M-token context, while the Dynamo serving recipe identifies 256K as the native model configuration and 1M as an explicitly enabled long-context serving profile. The model uses a hybrid Mamba-2, attention, and LatentMoE architecture with Multi-Token Prediction. It supports configurable reasoning through the chat template and is described as text-only. Available checkpoints include NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 and NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16, as well as a base BF16 checkpoint. No official provider-hosted per-token input or output price for the exact open-weight model was verified. Editorial scores reflect strong frontier reasoning and coding performance, moderate serving speed relative to smaller models, and substantial hardware costs.

Model guide

NVIDIA Nemotron 3 Ultra 550B-A55B: Frontier Reasoning for Long-Context Agents

NVIDIA Nemotron 3 Ultra 550B-A55B is a frontier-scale open-weight, text-only reasoning model with 550 billion total parameters and up to 55 billion active parameters per token. Its hybrid Mamba-2, attention, and LatentMoE architecture is designed for complex reasoning, coding, long-context analysis, retrieval-augmented generation, tool use, and multi-agent workflows. It supports configurable reasoning, native tool-use workflows through supported serving interfaces, and up to 1 million tokens through qualified long-context deployments, but its substantial hardware requirements make it better suited to enterprise and high-end infrastructure than to small or low-cost deployments.

What is NVIDIA Nemotron 3 Ultra 550B-A55B?

NVIDIA Nemotron 3 Ultra 550B-A55B is a frontier-scale large language model for text-based reasoning and generation. NVIDIA released it on June 4, 2026, as the largest member of the Nemotron 3 family. The model is available as open-weight checkpoints rather than only as a closed, provider-hosted chatbot. That gives organizations the option to deploy it through supported NVIDIA infrastructure or compatible third-party services, subject to the model's license, hardware, and operational requirements.

The name describes its scale: the model contains 550 billion total parameters, with up to 55 billion active parameters per token. In a sparse model, not every parameter is used for every token. This mixture-of-experts approach can provide a model with very large overall capacity without requiring all 550 billion parameters to participate in each individual calculation. It does not, however, make the model small or inexpensive to operate. Serving the model still requires high-memory, multi-GPU infrastructure.

Nemotron 3 Ultra is aimed at difficult tasks rather than casual chat. Its intended uses include multi-step reasoning, coding, research and synthesis, long-document analysis, high-stakes retrieval-augmented generation, tool-using agents, and multi-agent enterprise systems.

Architecture and reasoning capabilities

The model uses a hybrid architecture that combines Mamba-2 layers, selected attention layers, latent mixture-of-experts layers, and Multi-Token Prediction components. Mamba-2 is a state-space sequence architecture intended to process sequences efficiently, while attention layers help the model handle relationships between specific parts of the input. LatentMoE routes work among specialized expert components, and Multi-Token Prediction can support inference acceleration in compatible deployments.

For users, the most relevant feature is configurable reasoning. Supported chat templates can enable or disable extended reasoning, allowing a deployment to choose between more deliberate responses and a potentially more economical or lower-latency mode. The model is therefore suitable for tasks where intermediate reasoning is useful, such as decomposing a software problem, comparing evidence across long documents, planning a sequence of tool calls, or synthesizing information retrieved from an enterprise knowledge base.

The reasoning and coding ratings in the accompanying model data are editorial evaluations, not NVIDIA-published scores. They reflect the model's intended frontier reasoning and software-engineering positioning, but they should not be read as a substitute for testing the exact checkpoint, precision, serving stack, and prompts used in a production system.

Context window and long-context deployment

NVIDIA describes Nemotron 3 Ultra as supporting up to 1 million tokens of context. A context window is the amount of input and generated conversation history that a deployment can process within one request or session. A million-token context can be useful for large document collections, extended codebases, long research records, and agent histories that would otherwise need to be divided into many smaller requests.

There is an important deployment distinction. NVIDIA's Dynamo recipe identifies 256K tokens as the native model configuration, while qualified long-context serving profiles can override the serving framework's model-length guardrail to enable 1M-token operation. The 1M figure should therefore be understood as a supported long-context serving capability, not necessarily as the default setting in every installation.

The supplied research does not verify a separate maximum output-token limit for this exact model. Output capacity will depend on the serving configuration and the remaining space within the selected context limit. Operators should confirm the limits of the particular NIM, Dynamo, or third-party deployment rather than assuming that the full context window is available exclusively for generated output.

Supported modalities and tool use

Nemotron 3 Ultra is text-only. It accepts text input and produces text output; the supplied specifications do not identify native image, audio, or video input or output. It should not be selected when the model itself must interpret images, transcribe audio, or generate media. A separate multimodal system or preprocessing pipeline would be required for those workflows, and no such combination is established by the supplied model specifications.

The model supports tool calling and agentic workflows through supported NVIDIA serving interfaces. In practical terms, an application can provide descriptions of external functions or tools, allow the model to decide when a tool is relevant, execute that tool outside the model, and return the result for another reasoning step. The model does not independently access the web or enterprise systems simply because it supports tool use. Those connections must be implemented and permissioned by the application or serving environment.

Tool support makes the model relevant to agents that retrieve records, call business systems, run approved code utilities, or coordinate several specialized steps. It also introduces operational responsibilities: applications must validate arguments, control permissions, handle tool failures, and prevent untrusted retrieved content from being treated as an instruction.

Deployment, speed, and hardware trade-offs

NVIDIA documents deployments using high-memory Blackwell or Hopper GPUs. Supported configurations include multi-GPU systems based on H100, H200, B200, B300, GB200, and related platforms, depending on the selected precision and serving profile. BF16 and NVFP4 checkpoints provide different deployment trade-offs. BF16 generally preserves a higher-precision representation, while NVFP4 can reduce the memory and compute requirements of inference when the deployment supports it.

The model's sparse architecture and Multi-Token Prediction features may improve efficiency compared with a dense model of the same total parameter count, but Nemotron 3 Ultra remains a very large model. Its practical speed depends on GPU type, number of GPUs, precision, batching, context length, concurrency, and serving software. Long prompts and extended reasoning can also increase latency and resource consumption.

The editorial speed score is moderate relative to smaller models, while the editorial cost score reflects substantial infrastructure requirements. These are subjective evaluations based on the model's scale and deployment profile, not published NVIDIA service-level guarantees. Organizations should benchmark representative prompts before committing to a production architecture.

Pricing and access

No official provider-hosted per-token input or output price for the exact Nemotron 3 Ultra 550B-A55B checkpoint was verified in the supplied research. The model is open-weight, so its cost structure is different from that of a conventional hosted API model. A self-hosted deployment may involve GPU acquisition or rental, storage, networking, power, orchestration, monitoring, support, and licensing considerations. Access through an NVIDIA service or a third-party inference provider may have its own pricing and availability terms.

Because no verified input or output price is available for this exact model, a per-token comparison with smaller hosted models would be misleading. The relevant economic question is whether the model's quality and long-context or agentic capabilities justify the infrastructure required for the workload. Lower-precision NVFP4 serving may improve the hardware equation, but it does not turn the model into a lightweight deployment.

Languages and practical use cases

NVIDIA identifies support for English, French, Spanish, Italian, German, Japanese, Korean, Hindi, Brazilian Portuguese, and Chinese. The model is particularly suited to workflows where a system must reason over substantial textual evidence or coordinate multiple steps.

  • Complex reasoning: Break down technical or business problems, compare alternatives, and produce a reasoned answer.
  • Software engineering: Analyze code, propose implementations, explain failures, and assist with multi-step development tasks.
  • Long-document analysis: Review large collections of contracts, research materials, policies, or technical documentation when the deployment is configured for the required context length.
  • Enterprise RAG: Combine retrieved organizational information with extended reasoning, while keeping retrieval, access control, and source validation in the surrounding application.
  • Tool-using agents: Plan and execute workflows that call approved functions or enterprise services.
  • Multi-agent systems: Act as a high-capability reasoning component in systems that divide work among several specialized agents.

Main limitations

The first limitation is operational scale. This is not a practical default choice for a small local deployment, a modest developer workstation, or a low-latency consumer assistant. Its multi-GPU requirements and infrastructure costs can outweigh its benefits when a smaller model can complete the task adequately.

The second limitation is modality. The model does not natively handle images, audio, or video according to the supplied specifications. It is also not documented here as having a built-in web-search capability. External information retrieval and media processing must be supplied by surrounding systems.

The third limitation is the difference between advertised maximum context and default configuration. Although NVIDIA documents up to 1M tokens, the native Dynamo configuration is 256K and the longer profile requires an explicitly supported deployment. Long contexts can also increase memory use, latency, and cost.

Finally, open weights do not eliminate deployment complexity. Teams remain responsible for infrastructure, model serving, security, monitoring, prompt and tool policies, and validating the model's behavior for their domain. The supplied research also does not verify a fixed maximum output-token value, a universal JSON mode, caching, or batch API support for this exact checkpoint; those fields should remain deployment-specific rather than inferred.

When to choose Nemotron 3 Ultra

Choose Nemotron 3 Ultra when the workload genuinely benefits from frontier-scale text reasoning, long-context processing, coding ability, configurable reasoning, or tool-using agents, and when the organization can support high-end multi-GPU infrastructure. It is a strong candidate for enterprise research systems, complex software engineering assistants, long-context RAG, and agentic workflows where response quality matters more than minimal serving cost or latency.

A smaller reasoning or coding model is likely more appropriate for routine classification, short customer-support exchanges, simple extraction, high-volume batch processing, or applications with strict latency and budget targets. A model with native vision, speech, or media generation is more appropriate when multimodal input or output is central. Nemotron 3 Ultra's advantage is not universal feature coverage; it is the combination of very large model capacity, advanced reasoning orientation, open-weight deployment, and supported long-context serving.

Bottom line

NVIDIA Nemotron 3 Ultra 550B-A55B is a high-end open-weight text model designed for difficult reasoning and agentic workloads. Its 550B total parameters, 55B active parameters, hybrid architecture, configurable reasoning, tool-use support, and qualified 1M-token serving profile make it suitable for demanding enterprise and research deployments. Those capabilities come with substantial hardware and operational costs, no verified official per-token price for the exact checkpoint, and no native image, audio, or video capability. It is best evaluated as an infrastructure-level model for specialized workloads rather than as a general-purpose consumer chatbot.


Answers to Frequently Asked Questions

Does NVIDIA Nemotron 3 Ultra have an official per-token API price?
No verified official per-token input or output price was available for the exact NVIDIA Nemotron 3 Ultra 550B-A55B checkpoint. Because it is open-weight, costs may instead include GPU purchase or rental, storage, networking, power, serving, monitoring, support, licensing, or third-party provider fees.
Can NVIDIA Nemotron 3 Ultra process images, audio, or video?
No. Nemotron 3 Ultra is specified as a text-only model that accepts text input and produces text output. Image, audio, or video processing requires a separate multimodal model or an external preprocessing pipeline.
What hardware is required to run NVIDIA Nemotron 3 Ultra 550B-A55B?
Nemotron 3 Ultra requires high-memory, multi-GPU infrastructure. NVIDIA documents deployments using platforms such as H100, H200, B200, B300, and GB200, depending on the precision and serving profile. BF16 and NVFP4 offer different memory and performance trade-offs, but the model remains demanding to operate.
Does NVIDIA Nemotron 3 Ultra support a 1 million-token context window?
Yes, NVIDIA describes Nemotron 3 Ultra as supporting up to 1 million tokens, but this requires a qualified long-context serving configuration. NVIDIA's Dynamo recipe identifies 256K tokens as the native model configuration, so 1M-token operation should not be assumed to be enabled by default.
What is NVIDIA Nemotron 3 Ultra 550B-A55B?
NVIDIA Nemotron 3 Ultra 550B-A55B is an open-weight, frontier-scale language model designed for text reasoning, coding, long-document analysis, retrieval-augmented generation, tool-using agents, and multi-agent enterprise systems. It has 550 billion total parameters and up to 55 billion active parameters per token.


Sources 6
Provider

About NVIDIA AI