Nemotron Content Safety

Nemotron 3.5 Content Safety

by NVIDIA AI · Current; open-weight model and downloadable NVIDIA NIM; latest NGC container version bf16-v1.1 as of September 17, 2026

A 4B Gemma-based NVIDIA moderation model that evaluates multilingual text, one optional image, and AI responses for unsafe content, with standard taxonomy and custom-policy reasoning support.

Text Reasoning Coding
NVIDIA Nemotron 3.5 Content Safety is designed for a focused job: checking whether user prompts, images, and AI-generated responses violate safety requirements. It is not a general-purpose chatbot or content-generation model. The open-weight checkpoint supports 12 languages, accepts text with one optional image, and returns text-based safety labels that applications can use in moderation pipelines.
Outputs

What Nemotron 3.5 Content Safety can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

5/10 Reasoning
2/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Nemotron Content Safety
Model type Other
Context window 128K tokens
Release date 2026-03-16
Status Current; open-weight model and downloadable NVIDIA NIM; latest NGC container version bf16-v1.1 as of September 17, 2026
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was published in the reviewed NVIDIA or Hugging Face model documentation. The model is a moderation checkpoint rather than a general-purpose knowledge model, so its training-data dates should not be substituted for a knowledge cutoff.

Model notes

NVIDIA describes the model as a 4B decoder-only Transformer based on Gemma 3 4B IT, with a SigLIP vision encoder and LoRA fine-tuning whose weights were merged into the base model. It accepts text and one optional image, can optionally evaluate an assistant response, and supports standard safety categories plus custom-policy reasoning mode. The model card lists 12 supported languages and a context length of up to 128K tokens. Public deployment options include Hugging Face, NVIDIA NIM, Transformers, vLLM, and SGLang. Release dates differ by distribution surface: NVIDIA lists March 16, 2026 for Hugging Face and NGC and June 2, 2026 for Build.NVIDIA.com. The latest NGC model version identified in the current version history is bf16-v1.1, updated September 17, 2026. No official per-token hosted price was identified; the NIM page labels the endpoint downloadable and free. The model's text output is a classification string rather than a documented JSON-schema or legacy JSON-mode response.

Model guide

NVIDIA Nemotron 3.5 Content Safety: Multilingual Multimodal Moderation for AI Guardrails

NVIDIA Nemotron 3.5 Content Safety is an open-weight 4-billion-parameter moderation model for evaluating multilingual text, one optional image, and optional AI-generated responses. Based on Gemma 3 4B IT with a SigLIP vision encoder, it supports standard safety-taxonomy classification and custom-policy reasoning workflows for LLM and VLM guardrails.

What NVIDIA Nemotron 3.5 Content Safety is

Nemotron 3.5 Content Safety is a 4-billion-parameter model from NVIDIA that classifies potentially unsafe content. In simple terms, it acts as a safety reviewer placed before or after another AI model: it can inspect a user's request, an optional image, and an optional assistant response, then produce safety labels for the application to interpret.

The model is based on Google's Gemma 3 4B IT and was fine-tuned for multilingual, multimodal moderation and policy reasoning. Its specialization is important. This is not a general assistant intended to answer questions, write software, browse the web, or generate images. Its value comes from making safety decisions within an application-specific guardrail workflow.

NVIDIA provides the model as an open-weight downloadable checkpoint and through NVIDIA NIM. The reviewed NVIDIA and Hugging Face materials identify public deployment options including Hugging Face, Transformers, vLLM, and SGLang. NIM makes the model available as a downloadable endpoint, but the supplied research does not establish a separate hosted per-token pricing schedule.

How its moderation workflow works

A moderation request can contain a user prompt, one image, and an optional model response. This allows an application to assess the request and the response in the same interaction instead of treating input moderation and output moderation as completely separate tasks.

In its standard mode, the model returns text containing a required User Safety label and, when a response is supplied, an optional Response Safety label. It can also identify violated safety categories when the relevant content is unsafe. These are classification results intended for application-level parsing and enforcement, not polished explanations for end users.

The model also supports a custom-policy mode. Developers can provide domain-specific safety rules, after which the model can produce a concise reasoning trace followed by user and response safety labels. This may help teams audit why a moderation decision was made, but the trace should not be treated as a guaranteed explanation or as a replacement for human review in high-impact situations.

Supported inputs, outputs, and languages

Nemotron 3.5 Content Safety supports text input and one optional image. The image and text can be evaluated together, which is useful for cases such as a user caption paired with an uploaded image or a prompt asking another model to interpret visual material.

  • Text input: Supported for prompts and optional assistant responses.
  • Image input: One optional image can accompany the text request.
  • Audio and video input: Not identified as supported by the supplied model research.
  • Output: Text safety labels, optional violated categories, and optional custom-policy reasoning traces.
  • Generated media: The model does not generate images, audio, or video.

The model card lists English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Dutch, Italian, Korean, and Chinese. Multilingual support makes it more suitable for international moderation than a model trained primarily for English, although real-world quality can still vary by language, cultural context, and the type of content being reviewed.

Architecture, context, and deployment

The model uses a decoder-only Transformer architecture with Gemma 3 4B IT as its base. Its vision encoder is SigLIP, and NVIDIA documents square image processing at 896 by 896 pixels. The reported context length is up to 128K tokens, providing room for long prompts, policy instructions, and response text within the model's stated context window.

The public checkpoint is distributed in BF16 format. NVIDIA documents deployment on Linux systems with NVIDIA GPU hardware, including A100, H100, and RTX PRO 6000 Blackwell Server Edition systems. The model can also be deployed with common inference runtimes such as Transformers, vLLM, and SGLang, or through NVIDIA NIM.

The supplied specifications do not identify a maximum output-token limit. They also do not document a JSON-schema response mode. Although the model produces labels that an application may parse, its documented output is a classification string rather than a native structured-output contract. Production systems should therefore validate the response and handle malformed or unexpected text.

Pricing and cost considerations

No official per-token hosted price was identified in the supplied research. NVIDIA's NIM listing describes the endpoint as downloadable and free, while the open-weight model can be downloaded and run using compatible infrastructure. That does not mean deployment has no cost: users may still incur expenses for GPUs, cloud instances, storage, operations, monitoring, and support.

Its 4-billion-parameter size is a practical compromise. Compared with much larger general-purpose multimodal models, a specialized 4B moderator may be less expensive and faster to run for high-volume classification. However, the actual cost and latency depend on the selected GPU, runtime, batching configuration, image processing, prompt length, and deployment design. The research does not provide benchmark figures, so claims about exact throughput should be tested on the target hardware.

Main strengths and limitations

Where the model is strongest

  • Focused moderation: Its training and output format are centered on safety classification rather than general conversation.
  • Multimodal review: It can combine text and one image in a moderation decision.
  • Input and output guardrails: Applications can inspect both a user's request and an assistant's response.
  • Multilingual coverage: The model card lists 12 supported languages.
  • Custom policies: Developers can provide their own moderation rules for domain-specific workflows.
  • Deployment flexibility: The checkpoint is available for self-managed runtimes as well as NVIDIA NIM.
  • Long context: The documented context length is up to 128K tokens, which can accommodate lengthy policies and conversations.

Important limitations

The model is not a general-purpose assistant. It is not intended for open-ended chat, software development, web search, tool calling, image generation, audio processing, or video processing. Its coding capability is therefore not a meaningful use case, and the supplied research does not document function-calling support.

Its output is text classification rather than documented JSON-schema data. Teams building automated moderation should add a parser, validation rules, fallback behavior, logging, and escalation paths rather than assuming every response will match an ideal format.

Moderation decisions can be affected by language, cultural context, ambiguous wording, image quality, and the custom policy supplied by the developer. A model's classification should not be treated as infallible, particularly for high-impact decisions involving employment, access, safety, or legal status. Representative testing and human review may be necessary.

The model also requires suitable NVIDIA GPU infrastructure for the documented deployment environments. While its parameter count may support lower operating costs than much larger models, it is not automatically inexpensive or convenient for every team. Organizations without GPU operations experience may prefer a managed moderation service or another hosted solution.

Reasoning, coding, and tool-use profile

Nemotron 3.5 Content Safety has a narrow form of reasoning capability: custom-policy mode can produce a concise reasoning trace before its safety labels. This is useful for auditing policy application, but it should not be confused with broad problem-solving or extended chain-of-thought performance.

Coding is outside the model's purpose. It may classify text about code or software if that text is part of a moderation request, but it is not a code-generation model. Similarly, the model does not independently browse the web, call external tools, retrieve live information, or take actions. Any workflow involving those functions must provide them through the surrounding application, not through a documented native capability of this checkpoint.

When to choose this model

Nemotron 3.5 Content Safety is a reasonable choice when the primary requirement is self-managed moderation for multilingual text and optional images. It is especially relevant for:

  • Guarding prompts sent to large language or vision-language models.
  • Checking generated assistant responses before they reach users.
  • Moderating user-uploaded text and a single accompanying image.
  • Applying a custom safety policy instead of relying only on a general taxonomy.
  • Deploying an open-weight moderator in an NVIDIA GPU environment.
  • Keeping moderation close to an application's infrastructure rather than depending entirely on an external hosted API.

Its focused design can be preferable to sending every request to a larger general-purpose model. A smaller specialist may offer a better speed-and-cost trade-off for routine classification, particularly at scale, although this should be verified with application-specific tests rather than assumed from parameter count alone.

When another option may be more appropriate

A different model or service may be better if you need a conversational assistant, reliable structured JSON output, code generation, web access, audio or video moderation, multiple-image handling, or a managed API with clearly published usage prices and service-level guarantees. A larger multimodal model may also be more suitable when moderation depends on complex visual interpretation or nuanced reasoning beyond the safety categories and policies tested by this checkpoint.

Conversely, a simpler text-only classifier may be more efficient when images are never submitted. The choice should reflect the actual moderation surface: Nemotron 3.5 Content Safety is most compelling when multilingual text, image-aware moderation, custom rules, and self-managed NVIDIA deployment matter more than broad assistant capabilities.

Licensing and production use

The model is governed by the NVIDIA Open Model License Agreement, version 1.1, together with the Gemma Terms of Use and Gemma Prohibited Use Policy. NVIDIA describes it as ready for commercial use, but organizations should review the complete license terms and confirm that their planned deployment, data handling, and industry requirements are permitted.

Before production use, teams should test all supported languages and relevant image types, measure false positives and false negatives, verify how custom policies behave, and protect sensitive moderation inputs. They should also design a response-validation layer and define what happens when the model is uncertain, unavailable, or produces an output that cannot be parsed safely.


Answers to Frequently Asked Questions

How can NVIDIA Nemotron 3.5 Content Safety be deployed, and what does it cost?
The open-weight BF16 checkpoint can be deployed on Linux systems with NVIDIA GPUs, including A100, H100, and RTX PRO 6000 Blackwell Server Edition systems. Supported deployment options include Transformers, vLLM, SGLang, and NVIDIA NIM. No official per-token hosted price was identified; users may still incur costs for GPUs, cloud infrastructure, storage, operations, and support.
Can developers use custom safety policies with NVIDIA Nemotron 3.5 Content Safety?
Yes. Custom-policy mode allows developers to provide domain-specific safety rules. The model can then produce a concise reasoning trace followed by user and response safety labels, although the trace should not replace human review in high-impact decisions.
Can NVIDIA Nemotron 3.5 Content Safety moderate both user prompts and AI responses?
Yes. A moderation request can include a user prompt and an optional assistant response. The model can return a User Safety label and, when a response is supplied, a Response Safety label, along with violated safety categories when applicable.
What is NVIDIA Nemotron 3.5 Content Safety?
NVIDIA Nemotron 3.5 Content Safety is a 4-billion-parameter model designed to classify potentially unsafe content. It can review a user prompt, one optional image, and an optional assistant response, producing safety labels for application-level guardrails.
What languages and input types does NVIDIA Nemotron 3.5 Content Safety support?
The model supports text and one optional image in the same moderation request. Its model card lists English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Dutch, Italian, Korean, and Chinese. Audio and video inputs are not identified as supported.


Sources 4
Provider

About NVIDIA AI