Nemotron 3.5

NVIDIA Nemotron 3.5 Lightning 30B-A3B

by NVIDIA AI · Current open-weight model; NIM container available for early-access evaluation

Open-weight NVIDIA language model with a hybrid Mamba-2, MoE, and attention architecture. It targets fast, high-volume agent execution and supports reasoning, coding, long-context workloads, customization, and deployment through NVIDIA NIM, vLLM, and SGLang.

Text Reasoning Coding
NVIDIA Nemotron 3.5 Lightning 30B-A3B is designed for applications that need sustained text generation without paying the computational cost of activating all 30 billion of its parameters on every token. Its hybrid Mamba-2 and mixture-of-experts architecture uses approximately 3 billion active parameters per token, while selected attention layers help support long-context processing. NVIDIA provides BF16 reference weights and an NVFP4 quantized release, with deployment options including vLLM, SGLang, and NVIDIA NIM. The model supports reasoning mode, coding assistance, retrieval-augmented generation, chat, and agentic workflows, but it is a text-only model rather than a native image, audio, or video system.
Outputs

What NVIDIA Nemotron 3.5 Lightning 30B-A3B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Nemotron 3.5
Model type Reasoning
Context window 1M tokens
Knowledge cutoff 2025-09 pre-training; 2026-05 post-training
Release date 2026-08-11
Status Current open-weight model; NIM container available for early-access evaluation
Knowledge cutoff notes

The NVIDIA model card distinguishes between the pre-training data cutoff of September 2025 and the post-training data cutoff of May 2026. These are data freshness dates rather than a guarantee that every response is limited to one single cutoff.

Model notes

The canonical model is available in BF16 reference-weight and NVFP4 quantized releases. It has 30 billion total parameters and approximately 3 billion active parameters per token, using a hybrid architecture with interleaved Mamba-2 and MoE layers, selected attention layers, and multi-token prediction. NVIDIA documents reasoning mode through the chat template and provides inference integrations with vLLM, SGLang, and NVIDIA NIM. NVIDIA identifies the model for general-purpose reasoning and chat, AI agents, RAG pipelines, chatbots, coding assistance, and instruction following in English, coding languages, Spanish, French, German, Italian, and Japanese. The BF16 model card states that pre-training data has a September 2025 cutoff and post-training data has a May 2026 cutoff. NVIDIA also provides DSpark and DFlash speculative-decoding methods. Pricing is not populated because the official sources reviewed describe downloadable weights and NIM deployment rather than a universal per-token hosted price. Editorial scores are comparative estimates, not vendor ratings.

Model guide

NVIDIA Nemotron 3.5 Lightning: Fast Open-Weight Model for Long-Running AI Agents

NVIDIA Nemotron 3.5 Lightning 30B-A3B is an open-weight hybrid mixture-of-experts language model with 30 billion total parameters and approximately 3 billion active parameters per token. It combines Mamba-2, selected attention layers, mixture-of-experts routing, and multi-token prediction for efficient text generation in agent execution, reasoning, coding, retrieval-augmented generation, chat, and long-context workloads.

What is NVIDIA Nemotron 3.5 Lightning 30B-A3B?

NVIDIA Nemotron 3.5 Lightning 30B-A3B is an open-weight language model from NVIDIA's Nemotron family. It is intended for developers and organizations that want to run, customize, or integrate a capable text model in their own infrastructure rather than relying exclusively on a managed, per-token API.

The model has 30 billion total parameters, but approximately 3 billion are active for each token. Parameters are the learned numerical values that store patterns acquired during training. Activating only a portion of the model for each token can reduce the computation required for generation while retaining the capacity of a larger model. This is the main practical idea behind the model's mixture-of-experts, or MoE, design.

NVIDIA positions Nemotron 3.5 Lightning for high-volume and long-running agent execution, reasoning, coding assistance, chat, retrieval-augmented generation (RAG), and instruction following. The model is available in BF16 reference-weight and NVFP4 quantized variants. NVIDIA also provides a NIM container for deployment and documents integrations with vLLM and SGLang.

Architecture and context length

The model uses a hybrid architecture rather than relying on one mechanism throughout the network. It combines Mamba-2 layers, mixture-of-experts layers, selected attention layers, and multi-token prediction. Mamba-2 is a state-space architecture designed to process sequences efficiently, while attention layers can help the model make targeted relationships between tokens. The MoE layers route work through selected expert components instead of activating every parameter for every token.

This design is intended to balance long-context handling, generation speed, and model capacity. NVIDIA documents context support of up to 1,000,000 tokens in validated deployments. A context window includes the material supplied to the model and the conversation or generated content retained for the request. The one-million-token figure is therefore a deployment capability, not a guarantee that every hardware configuration or serving stack will handle that length at the same speed or memory cost.

The supplied specifications do not provide a maximum output-token limit. That limit can depend on the serving configuration, available memory, and the selected deployment tool, so applications should verify the limit in their chosen runtime instead of assuming that the full context window is available for output.

Reasoning, coding, and agent use

Nemotron 3.5 Lightning supports a reasoning mode through its chat template. In practical terms, this makes it suitable for tasks that benefit from intermediate problem solving, such as multi-step analysis, planning, code generation, and structured decision-making. Reasoning mode should not be interpreted as a guarantee of correctness: applications still need validation, especially when generated code or agent actions affect external systems.

Coding is one of the model's intended uses. NVIDIA identifies coding assistance and instruction following among its target workloads, and the model supports several programming languages in addition to English. The documented language coverage includes Spanish, French, German, Italian, and Japanese, as well as coding languages. This makes it a candidate for code explanation, generation, transformation, debugging assistance, and software-oriented agents.

The model is also suited to RAG systems. A RAG application retrieves relevant documents and places them in the model's input so that responses can be grounded in a private or specialized information source. Its long-context capability can be useful when an application needs to process large collections of retrieved material, lengthy task histories, or extended agent traces. However, a large context window does not remove the need for careful retrieval, chunking, source attribution, and prompt design.

Nemotron 3.5 Lightning is documented for agentic execution and tool use. It can serve as the language component of an agent that calls search, database, software, or business-system tools. The model itself is text-only, so tool invocation and action execution depend on the surrounding application or serving framework. Developers should define explicit tool schemas, permissions, validation, and failure handling rather than allowing generated text to trigger unrestricted actions.

Speed, cost, and deployment trade-offs

The model's approximately 3 billion active parameters, hybrid architecture, and multi-token prediction are intended to improve generation efficiency. NVIDIA also provides DSpark and DFlash speculative-decoding methods, which are designed to accelerate generation by predicting multiple tokens or using a faster auxiliary process before verification. Actual performance will depend on the GPU, quantization format, batch size, sequence length, runtime, and serving configuration.

The NVFP4 variant can reduce the memory and compute requirements compared with a BF16 deployment, but quantization may involve quality or compatibility trade-offs for some workloads. BF16 is the more conventional reference-weight format and may be preferable when hardware capacity and numerical behavior are more important than minimizing resource use. These are deployment choices rather than separate model capabilities.

There is no universal per-token price established in the supplied NVIDIA sources. The model is available as downloadable weights, and NVIDIA lists NIM deployment for evaluation and serving, but the reviewed material does not provide a single primary hosted price that applies to all users. Infrastructure, GPU, licensing, support, and NIM terms can therefore determine the real cost. Organizations comparing it with a managed API should include the operational cost of hosting, monitoring, scaling, and maintaining the model.

Modalities and important limitations

Nemotron 3.5 Lightning is a text language model. It accepts text input and produces text output. The supplied model specifications do not establish native image, audio, or video input, and it does not directly generate images, audio, video, music, or speech. Multimodal applications would need to add separate models or preprocessing components around it.

It is also not presented as a managed first-party token-priced service with one universal endpoint and fixed limits for every customer. Users who want a ready-made cloud chatbot, built-in web search, or a fully managed application may find a hosted general-purpose service more convenient. Nemotron 3.5 Lightning is better suited to teams that can manage model deployment or use an NVIDIA-supported serving layer.

Open weights provide more control, but they also transfer responsibility to the operator. Hardware selection, GPU memory, runtime configuration, security, scaling, logging, model updates, and output evaluation remain deployment concerns. The model card identifies a September 2025 pre-training data cutoff and a May 2026 post-training data cutoff. These dates describe the freshness of the training data and do not mean that the model has live knowledge or web access. The model has no built-in web search capability according to the supplied specifications.

When to choose Nemotron 3.5 Lightning

This model is a strong candidate when the priority is efficient, self-managed text generation for sustained workloads. Appropriate examples include:

  • Long-running agents that repeatedly plan, call tools, and summarize results.
  • High-volume coding assistance or software-development workflows.
  • RAG systems that need to process large retrieved contexts.
  • Private or specialized deployments where open weights and customization are important.
  • Reasoning and instruction-following applications that can benefit from a relatively small active parameter count.
  • Teams evaluating BF16 or NVFP4 deployment through vLLM, SGLang, or NVIDIA NIM.

Another type of model may be more appropriate when native vision, speech, image generation, or video generation is required. A managed API may also be preferable when the team does not want to operate GPUs or maintain an inference stack. Conversely, a much larger dense or expert model may be a better choice for tasks where maximum capability matters more than throughput and infrastructure cost. Nemotron 3.5 Lightning's value is its balance: it targets practical speed and operating efficiency while retaining a 30-billion-parameter model capacity.

Verified specifications at a glance

SpecificationDetails
ProviderNVIDIA
Model familyNemotron 3.5
Model typeOpen-weight hybrid mixture-of-experts language model
Parameters30B total; approximately 3B active per token
ArchitectureHybrid Mamba-2, MoE, selected attention layers, and multi-token prediction
ContextUp to 1,000,000 tokens in supported or validated deployments
Input and outputText input and text output
ReasoningReasoning mode supported through the chat template
FormatsBF16 reference weights and NVFP4 quantized release
DeploymentvLLM, SGLang, NVIDIA NIM, and related NVIDIA tooling
PricingNo universal hosted per-token price verified in the supplied sources
Maximum outputNot specified in the supplied research

Overall, NVIDIA Nemotron 3.5 Lightning 30B-A3B is aimed at developers who need an efficient open-weight model for text-heavy agents, reasoning, coding, and long-context applications. Its active-parameter design and deployment flexibility are its clearest advantages, while the need for infrastructure and its lack of native non-text modalities are the main reasons to choose another option.


Answers to Frequently Asked Questions

Does Nemotron 3.5 Lightning support images, audio, video, or built-in web search?
No. Nemotron 3.5 Lightning is a text-only model that accepts text input and produces text output. The supplied specifications do not establish native image, audio, or video capabilities, and the model does not have built-in web search. Multimodal or web-enabled applications require additional models or external services.
Is NVIDIA Nemotron 3.5 Lightning suitable for coding and AI agents?
Yes. The model is intended for coding assistance, reasoning, planning, tool use, and long-running agentic workflows. It can generate and transform code, assist with debugging, process retrieved information, and serve as the language component of an agent, but external applications must provide tool execution, permissions, validation, and failure handling.
How many tokens can NVIDIA Nemotron 3.5 Lightning handle?
NVIDIA documents support for context lengths of up to 1,000,000 tokens in validated deployments. The actual usable context length, speed, memory cost, and maximum output length depend on the hardware, serving framework, and deployment configuration.
What are the main deployment formats and runtimes for Nemotron 3.5 Lightning?
Nemotron 3.5 Lightning is available in BF16 reference-weight and NVFP4 quantized variants. NVIDIA documents deployment through NVIDIA NIM, vLLM, and SGLang, with actual performance depending on the GPU, quantization format, batch size, sequence length, and serving configuration.
What is NVIDIA Nemotron 3.5 Lightning 30B-A3B?
NVIDIA Nemotron 3.5 Lightning 30B-A3B is an open-weight hybrid mixture-of-experts language model designed for self-managed text generation, long-running AI agents, reasoning, coding, retrieval-augmented generation, and instruction following. It has 30 billion total parameters, with approximately 3 billion active for each token.


Sources 5
Provider

About NVIDIA AI