What is NVIDIA Nemotron 3.5 Lightning 30B-A3B?
NVIDIA Nemotron 3.5 Lightning 30B-A3B is an open-weight language model from NVIDIA's Nemotron family. It is intended for developers and organizations that want to run, customize, or integrate a capable text model in their own infrastructure rather than relying exclusively on a managed, per-token API.
The model has 30 billion total parameters, but approximately 3 billion are active for each token. Parameters are the learned numerical values that store patterns acquired during training. Activating only a portion of the model for each token can reduce the computation required for generation while retaining the capacity of a larger model. This is the main practical idea behind the model's mixture-of-experts, or MoE, design.
NVIDIA positions Nemotron 3.5 Lightning for high-volume and long-running agent execution, reasoning, coding assistance, chat, retrieval-augmented generation (RAG), and instruction following. The model is available in BF16 reference-weight and NVFP4 quantized variants. NVIDIA also provides a NIM container for deployment and documents integrations with vLLM and SGLang.
Architecture and context length
The model uses a hybrid architecture rather than relying on one mechanism throughout the network. It combines Mamba-2 layers, mixture-of-experts layers, selected attention layers, and multi-token prediction. Mamba-2 is a state-space architecture designed to process sequences efficiently, while attention layers can help the model make targeted relationships between tokens. The MoE layers route work through selected expert components instead of activating every parameter for every token.
This design is intended to balance long-context handling, generation speed, and model capacity. NVIDIA documents context support of up to 1,000,000 tokens in validated deployments. A context window includes the material supplied to the model and the conversation or generated content retained for the request. The one-million-token figure is therefore a deployment capability, not a guarantee that every hardware configuration or serving stack will handle that length at the same speed or memory cost.
The supplied specifications do not provide a maximum output-token limit. That limit can depend on the serving configuration, available memory, and the selected deployment tool, so applications should verify the limit in their chosen runtime instead of assuming that the full context window is available for output.
Reasoning, coding, and agent use
Nemotron 3.5 Lightning supports a reasoning mode through its chat template. In practical terms, this makes it suitable for tasks that benefit from intermediate problem solving, such as multi-step analysis, planning, code generation, and structured decision-making. Reasoning mode should not be interpreted as a guarantee of correctness: applications still need validation, especially when generated code or agent actions affect external systems.
Coding is one of the model's intended uses. NVIDIA identifies coding assistance and instruction following among its target workloads, and the model supports several programming languages in addition to English. The documented language coverage includes Spanish, French, German, Italian, and Japanese, as well as coding languages. This makes it a candidate for code explanation, generation, transformation, debugging assistance, and software-oriented agents.
The model is also suited to RAG systems. A RAG application retrieves relevant documents and places them in the model's input so that responses can be grounded in a private or specialized information source. Its long-context capability can be useful when an application needs to process large collections of retrieved material, lengthy task histories, or extended agent traces. However, a large context window does not remove the need for careful retrieval, chunking, source attribution, and prompt design.
Nemotron 3.5 Lightning is documented for agentic execution and tool use. It can serve as the language component of an agent that calls search, database, software, or business-system tools. The model itself is text-only, so tool invocation and action execution depend on the surrounding application or serving framework. Developers should define explicit tool schemas, permissions, validation, and failure handling rather than allowing generated text to trigger unrestricted actions.
Speed, cost, and deployment trade-offs
The model's approximately 3 billion active parameters, hybrid architecture, and multi-token prediction are intended to improve generation efficiency. NVIDIA also provides DSpark and DFlash speculative-decoding methods, which are designed to accelerate generation by predicting multiple tokens or using a faster auxiliary process before verification. Actual performance will depend on the GPU, quantization format, batch size, sequence length, runtime, and serving configuration.
The NVFP4 variant can reduce the memory and compute requirements compared with a BF16 deployment, but quantization may involve quality or compatibility trade-offs for some workloads. BF16 is the more conventional reference-weight format and may be preferable when hardware capacity and numerical behavior are more important than minimizing resource use. These are deployment choices rather than separate model capabilities.
There is no universal per-token price established in the supplied NVIDIA sources. The model is available as downloadable weights, and NVIDIA lists NIM deployment for evaluation and serving, but the reviewed material does not provide a single primary hosted price that applies to all users. Infrastructure, GPU, licensing, support, and NIM terms can therefore determine the real cost. Organizations comparing it with a managed API should include the operational cost of hosting, monitoring, scaling, and maintaining the model.
Modalities and important limitations
Nemotron 3.5 Lightning is a text language model. It accepts text input and produces text output. The supplied model specifications do not establish native image, audio, or video input, and it does not directly generate images, audio, video, music, or speech. Multimodal applications would need to add separate models or preprocessing components around it.
It is also not presented as a managed first-party token-priced service with one universal endpoint and fixed limits for every customer. Users who want a ready-made cloud chatbot, built-in web search, or a fully managed application may find a hosted general-purpose service more convenient. Nemotron 3.5 Lightning is better suited to teams that can manage model deployment or use an NVIDIA-supported serving layer.
Open weights provide more control, but they also transfer responsibility to the operator. Hardware selection, GPU memory, runtime configuration, security, scaling, logging, model updates, and output evaluation remain deployment concerns. The model card identifies a September 2025 pre-training data cutoff and a May 2026 post-training data cutoff. These dates describe the freshness of the training data and do not mean that the model has live knowledge or web access. The model has no built-in web search capability according to the supplied specifications.
When to choose Nemotron 3.5 Lightning
This model is a strong candidate when the priority is efficient, self-managed text generation for sustained workloads. Appropriate examples include:
- Long-running agents that repeatedly plan, call tools, and summarize results.
- High-volume coding assistance or software-development workflows.
- RAG systems that need to process large retrieved contexts.
- Private or specialized deployments where open weights and customization are important.
- Reasoning and instruction-following applications that can benefit from a relatively small active parameter count.
- Teams evaluating BF16 or NVFP4 deployment through vLLM, SGLang, or NVIDIA NIM.
Another type of model may be more appropriate when native vision, speech, image generation, or video generation is required. A managed API may also be preferable when the team does not want to operate GPUs or maintain an inference stack. Conversely, a much larger dense or expert model may be a better choice for tasks where maximum capability matters more than throughput and infrastructure cost. Nemotron 3.5 Lightning's value is its balance: it targets practical speed and operating efficiency while retaining a 30-billion-parameter model capacity.
Verified specifications at a glance
| Specification | Details |
|---|---|
| Provider | NVIDIA |
| Model family | Nemotron 3.5 |
| Model type | Open-weight hybrid mixture-of-experts language model |
| Parameters | 30B total; approximately 3B active per token |
| Architecture | Hybrid Mamba-2, MoE, selected attention layers, and multi-token prediction |
| Context | Up to 1,000,000 tokens in supported or validated deployments |
| Input and output | Text input and text output |
| Reasoning | Reasoning mode supported through the chat template |
| Formats | BF16 reference weights and NVFP4 quantized release |
| Deployment | vLLM, SGLang, NVIDIA NIM, and related NVIDIA tooling |
| Pricing | No universal hosted per-token price verified in the supplied sources |
| Maximum output | Not specified in the supplied research |
Overall, NVIDIA Nemotron 3.5 Lightning 30B-A3B is aimed at developers who need an efficient open-weight model for text-heavy agents, reasoning, coding, and long-context applications. Its active-parameter design and deployment flexibility are its clearest advantages, while the need for infrastructure and its lack of native non-text modalities are the main reasons to choose another option.

