Nemotron 3 Nano Omni

NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning

by NVIDIA AI · Current; generally available open-weight checkpoint with hosted NVIDIA NIM access

Open-weight NVIDIA omnimodal reasoning model with 30B total and approximately 3B active parameters. It accepts text, images, video, and audio, supports long multimodal context, and is available in BF16, FP8, and NVFP4 checkpoints for self-hosted and NIM deployment.

Text Reasoning Coding
NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning combines a NemotronH hybrid Mamba-transformer language backbone with dedicated vision and audio encoders. It is designed for long-context multimodal understanding, document analysis, video and audio comprehension, reasoning, tool-oriented agent workflows, and self-hosted deployment.
Outputs

What NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Streaming Fine-tuning Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
7/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Nemotron 3 Nano Omni
Model type Multimodal
Context window 262K tokens
Maximum output 131K tokens
Release date 2026-04-27
Status Current; generally available open-weight checkpoint with hosted NVIDIA NIM access
Knowledge cutoff notes

NVIDIA's reviewed model and architecture documentation does not provide a direct, authoritative knowledge-cutoff date for this exact checkpoint.

Model notes

The canonical open-weight checkpoint is NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16. NVIDIA also publishes FP8 and NVFP4 variants; these are quantization variants of the same model rather than separate model entities. The architecture has 30B total parameters and approximately 3B active parameters. NVIDIA documentation describes the context as 262K tokens, while the research report describes it as 256K tokens. The model accepts text, images, video, and audio and returns text. NVIDIA documents tool-oriented reasoning and NIM deployment, but exact structured-output, JSON-mode, batch, and hosted pricing behavior is deployment-specific or not independently documented for the exact model. Editorial scores are comparative estimates, not NVIDIA specifications.

Model guide

NVIDIA Nemotron 3 Nano Omni: Open-Weight Reasoning Across Text, Image, Video and Audio

NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning is an open-weight omnimodal reasoning model that processes text, images, video, and audio through a unified 30B-parameter, 3B-active hybrid mixture-of-experts architecture.

What NVIDIA Nemotron 3 Nano Omni is

NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning is an open-weight model for reasoning over text, images, video, and audio. NVIDIA provides downloadable checkpoints and deployment paths through Hugging Face, Transformers, vLLM, NeMo, NVIDIA NIM, and the NVIDIA Build platform. It is part of NVIDIA's current Nemotron model lineup, but this specific model is focused on unified multimodal understanding rather than image, video, or speech generation.

The model has 30 billion total parameters, with approximately 3 billion active parameters used during each forward pass. This is a mixture-of-experts, or MoE, design: the complete model retains substantial capacity, while only selected expert components process each token. In practical terms, that can reduce the computation required per token compared with using all 30 billion parameters on every step, although actual speed and memory use depend on the checkpoint, hardware, quantization, serving engine, and input type.

The model is intended for applications such as document intelligence, OCR, chart interpretation, long-video analysis, audio understanding, cross-modal question answering, voice-oriented agents, and self-hosted enterprise inference.

Architecture and supported modalities

Nemotron 3 Nano Omni uses an encoder-projector-decoder architecture. Separate components first convert different kinds of media into representations that the language decoder can process. The decoder then performs the central reasoning and text-generation work across those inputs.

  • Text: Native text input and text generation.
  • Images: A C-RADIOv4-H vision encoder with dynamic-resolution processing.
  • Video: Temporal processing using 3D convolutions and Efficient Video Sampling.
  • Audio: A Parakeet-based audio encoder for speech and other audio understanding tasks.

This unified design is important because the model can reason across modalities in one workflow. For example, an application could provide a scanned document, its accompanying chart, and an audio explanation, then ask for a text answer that connects the information. The supplied research identifies text as the model's output modality; it does not natively produce images, video, music, or speech audio.

Context window and output limit

NVIDIA's documentation lists a maximum context length of 262,144 tokens. A related research report describes the model as having a 256K-token context, so the two figures should be understood as different ways of describing approximately the same long-context capability. NVIDIA reports progressive context training from 16K to 49K and then to 262K tokens.

The recorded maximum output limit is 131,072 tokens. In practice, the usable output length is constrained by the total context budget, the selected deployment configuration, available GPU memory, and the serving software. Large media inputs can also consume part of the available context, so a nominal 262K-token limit should not be interpreted as 262K text tokens available after every image, video, or audio input.

This context size makes the model suitable for long documents, extended transcripts, and substantial multimodal sequences. It does not remove the need to manage input quality: very long or poorly prepared media can still make retrieval, interpretation, and answer verification more difficult.

Reasoning, coding and tool-oriented workflows

Reasoning is a central purpose of this checkpoint. NVIDIA positions it for multimodal reasoning, document understanding, OCR, audio-visual reasoning, voice interaction, video comprehension, and agent-style workflows. The model can combine visual, audio, and textual evidence instead of treating each input as an isolated task.

The research data gives the model an editorial reasoning score of 8 out of 10. That score is a comparative assessment rather than an NVIDIA-published benchmark result and should not be confused with an official standardized ranking. NVIDIA's materials do report strong performance across several multimodal and text-reasoning areas, but the supplied research does not provide a complete set of independently comparable benchmark numbers.

Tool-oriented reasoning is documented, and the model is recorded as supporting tool use. This makes it a possible component in an agent pipeline where the model decides when to call an external function or service. Exact function-calling syntax, structured-output guarantees, and JSON behavior depend on the deployment surface. There is no verified universal JSON mode for this model.

The model can be used for code-related reasoning and tool workflows, and its editorial coding score is 7 out of 10. However, the supplied research does not establish a dedicated coding specialization or a published coding benchmark for this exact checkpoint. Teams choosing it primarily for software development should validate code-generation accuracy, repository-scale context handling, and tool-call reliability against a model designed and evaluated specifically for coding.

Speed, cost and deployment trade-offs

NVIDIA publishes BF16, FP8, and NVFP4 versions. BF16 is identified as the general-availability release and is intended for higher-precision deployment. FP8 and NVFP4 can reduce memory requirements on compatible hardware, but quantization may affect quality or supported operations depending on the workload.

NVIDIA's research report describes more than 500 output tokens per second at single-stream concurrency on an NVIDIA B200 with an optimized NVFP4 deployment. This is a provider-reported configuration-specific result, not a universal speed guarantee. Performance can change substantially with GPU type, batch size, media preprocessing, context length, quantization, and inference engine.

The editorial speed score is 8 out of 10 and the editorial cost score is also 8 out of 10. These are comparative estimates based on the model's active-parameter design and available quantized checkpoints, not provider specifications. The MoE architecture and quantization options can improve the capacity-to-compute trade-off, but self-hosting still requires substantial GPU resources, especially with BF16.

No standardized public token price for this exact model is established in the reviewed documentation. NVIDIA provides a hosted trial endpoint through Build, while downloadable checkpoints support self-hosting and NIM deployment. Operating cost therefore depends on whether the model is run on owned hardware, rented infrastructure, or a hosted NVIDIA service, as well as on the selected precision and usage level.

Main strengths and limitations

Strengths

  • Broad native input support: Text, images, video, and audio are handled within one model workflow.
  • Long context: The approximately 256K–262K-token context is useful for large documents, transcripts, and extended multimodal sequences.
  • Efficient MoE structure: Approximately 3B active parameters per forward pass can offer a more favorable compute trade-off than dense processing at the same total parameter count.
  • Open-weight deployment: BF16, FP8, and NVFP4 checkpoints support different accuracy, memory, and throughput requirements.
  • Enterprise-oriented flexibility: The model can be deployed through common open-model tooling, NVIDIA NIM, or self-hosted infrastructure under the NVIDIA Nemotron Open Model License, subject to its terms.

Limitations

  • Hardware requirements: Self-hosting, particularly with BF16, requires significant GPU memory and a compatible inference environment.
  • Serving complexity: Multimodal deployment requires media preprocessing, model-specific code, and an inference engine that supports the relevant inputs.
  • Configuration-dependent speed: The reported B200 throughput does not represent performance on ordinary workstations or every multimodal request.
  • Text-only output: The model understands several media types but does not natively generate images, video, music, or speech audio.
  • Unclear hosted economics: A universal public token price for the exact model is not established in the supplied documentation.
  • Deployment-specific API behavior: JSON mode, structured outputs, batching, caching, and exact tool-call behavior may vary or remain undocumented for a particular serving surface.

Best use cases

Nemotron 3 Nano Omni is a strong fit when one system must interpret several media types and return a reasoned text response. Suitable examples include extracting information from lengthy reports, answering questions about scanned forms and charts, reviewing video recordings, summarizing meetings with audio and visual context, and building assistants that connect speech, documents, and external tools.

It is also a practical candidate for organizations that need an open-weight model they can deploy on their own infrastructure. The available quantization variants allow teams to evaluate different memory and throughput trade-offs, while NIM and other NVIDIA deployment paths can reduce some of the integration work when the required hardware and software are available.

When to choose this model

Choose NVIDIA Nemotron 3 Nano Omni when multimodal understanding is more important than a simple text-only chatbot interface, when long context is valuable, or when self-hosting and open-weight access are important. It is particularly relevant for document intelligence, OCR, video and audio analysis, multimodal agents, and enterprise workflows built around NVIDIA hardware.

Another option may be more appropriate when the application needs native image, video, music, or speech generation; lightweight edge deployment; a clearly documented hosted token price; or guaranteed structured-output behavior. A text-specialized or coding-specialized model may also be easier to operate for workloads that do not need media understanding. Likewise, a smaller model may offer a better cost and latency profile when inputs are short and the task is limited to ordinary text generation.

Availability and license

The model is available through NVIDIA's Build platform and downloadable repositories on Hugging Face, including BF16, FP8, and NVFP4 variants. NVIDIA documents deployment with Transformers, vLLM, NeMo, and NIM. The model is released under the NVIDIA Nemotron Open Model License, which supports enterprise, commercial, and on-premises use subject to the license terms.

Overall, Nemotron 3 Nano Omni occupies a specific position: it is not simply a smaller text model, and it is not a generative media system. Its value comes from combining long-context reasoning, broad multimodal input support, open-weight deployment, and an MoE architecture that limits active computation relative to total capacity. Those benefits are most relevant when the application can justify the infrastructure and integration work required to serve a large multimodal model.


Answers to Frequently Asked Questions

What are the main limitations of NVIDIA Nemotron 3 Nano Omni?
The model requires substantial GPU resources, particularly in BF16, and multimodal serving can involve complex preprocessing and model-specific integration. Its speed depends on hardware, quantization, context length, media type, and inference engine. It also has text-only output, no universally verified JSON mode, and no standardized public token price established for the exact model.
How can NVIDIA Nemotron 3 Nano Omni be deployed?
The model is available through NVIDIA Build and downloadable Hugging Face repositories, with BF16, FP8, and NVFP4 variants. NVIDIA documents deployment using Transformers, vLLM, NeMo, and NVIDIA NIM. It can be self-hosted or operated through compatible NVIDIA infrastructure, subject to hardware, software, and license requirements.
How many parameters and how much context does Nemotron 3 Nano Omni have?
Nemotron 3 Nano Omni has 30 billion total parameters and approximately 3 billion active parameters per forward pass because it uses a mixture-of-experts architecture. NVIDIA lists a maximum context length of 262,144 tokens, while related research describes it as approximately 256K tokens. The recorded maximum output limit is 131,072 tokens.
What is NVIDIA Nemotron 3 Nano Omni?
NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning is an open-weight mixture-of-experts model designed to reason across text, images, video, and audio. It is intended for applications such as document intelligence, OCR, chart interpretation, long-video analysis, audio understanding, multimodal question answering, and enterprise inference.
Which modalities does NVIDIA Nemotron 3 Nano Omni support?
The model accepts text, images, video, and audio as inputs and produces text as its output. It can combine information from multiple modalities in one workflow, but it does not natively generate images, video, music, or speech audio.


Sources 8
Provider

About NVIDIA AI