Nemotron 3 VoiceChat

NVIDIA Nemotron 3 VoiceChat

by NVIDIA AI · Early access; available for evaluation through NVIDIA NIM and qualified access programs

A 12-billion-parameter NVIDIA model for full-duplex speech-to-speech interaction. Nemotron 3 VoiceChat combines streaming speech understanding, language generation, transcription, and text-to-speech for low-latency voice agents, with early-access availability through NVIDIA NIM and a documented focus on NVIDIA Hopper hardware.

Text Speech Reasoning Coding
NVIDIA Nemotron 3 VoiceChat is an early-access model designed for low-latency, interruptible voice conversations rather than general-purpose text generation. It accepts text prompts and streaming 16 kHz user speech, then returns agent text, a transcription of the user’s speech, and synthesized 22.05 kHz agent speech from a unified architecture.
Outputs

What NVIDIA Nemotron 3 VoiceChat can produce

Text Speech
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Streaming Fine-tuning Multimodal output
Model profile

Performance characteristics

5/10 Reasoning
3/10 Coding
9/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Nemotron 3 VoiceChat
Model type Multimodal
Release date 2026-03-16
Status Early access; available for evaluation through NVIDIA NIM and qualified access programs
Knowledge cutoff notes

NVIDIA does not publish a knowledge-cutoff date for the exact Nemotron 3 VoiceChat model. Its speech-and-language system is optimized for real-time conversational interaction rather than exposing a documented standalone language-model cutoff.

Model notes

The exact canonical model identity is NVIDIA Nemotron 3 VoiceChat, served in NVIDIA documentation as the nemotron-voicechat NIM. It has 12B parameters and uses a hybrid Mamba/Transformer architecture combining a Fast Conformer speech encoder, a Nemotron Nano v2 language backbone, and an NVIDIA TTS decoder and codec. It accepts text prompts and 16 kHz user audio, and returns agent text, user-speech transcription, and 22.05 kHz agent speech. The model is intended for early-access evaluation only and is optimized for NVIDIA Hopper hardware, especially H100 GPUs, with vLLM on Linux. NVIDIA documents a guided fine-tuning path through the early-access program, but does not publish a general context-window limit, maximum output-token limit, batch API, caching API, or separate JSON-mode capability for this exact model.

Cost

Model pricing

Input Free endpoint for NVIDIA NIM trial access; no general production price published
Output Free endpoint for NVIDIA NIM trial access; no general production price published
Model guide

NVIDIA Nemotron 3 VoiceChat: Full-Duplex Speech-to-Speech for Real-Time Voice Agents

NVIDIA Nemotron 3 VoiceChat is a 12-billion-parameter, full-duplex speech-to-speech model for real-time voice agents. It combines streaming speech understanding, language generation, transcription, and speech synthesis so users can interrupt conversations naturally while the system continues processing audio.

What is NVIDIA Nemotron 3 VoiceChat?

NVIDIA Nemotron 3 VoiceChat is a 12-billion-parameter, full-duplex speech-to-speech model. “Full-duplex” means that the system is built for simultaneous two-way audio interaction: the user can speak, the model can respond, and the user can interrupt instead of waiting for a conventional turn-based exchange to finish.

The model is intended for real-time voice-agent applications such as conversational interfaces, speech-to-speech research, and enterprise systems running on NVIDIA hardware. It is not positioned as a general-purpose chatbot for long written prompts, image analysis, coding, or broad multilingual communication.

NVIDIA provides the model through an early-access program and an NVIDIA NIM experience. NIM is NVIDIA’s model-serving software and hosted model ecosystem; in NVIDIA documentation, this model is served under the NIM identity nemotron-voicechat. Access and deployment requirements may differ between the hosted evaluation endpoint and self-managed infrastructure.

How the model processes speech

Nemotron 3 VoiceChat uses a hybrid Mamba/Transformer architecture. Its documented components include a Fast Conformer speech encoder, a Nemotron Nano v2 language backbone, and an NVIDIA text-to-speech decoder and codec. In practical terms, the speech encoder turns incoming audio into information the language system can process, while the decoder and codec produce the audio response.

The model accepts text prompts as well as streaming user audio sampled at 16 kHz. Its documented outputs include:

  • Agent-generated text.
  • A transcription of the user’s speech.
  • Agent speech sampled at 22.05 kHz.

This combination makes it different from a text-only language model connected to separate speech-recognition and text-to-speech services. A multi-service pipeline may still be preferable when an application needs independently replaceable components, broader language coverage, or more granular control over transcription and synthesis.

Modalities and capabilities

CapabilitySupported or documented behavior
Text inputYes; text prompts are supported.
Audio inputYes; streaming user speech at 16 kHz.
Text outputYes; the model returns agent text and user-speech transcription.
Audio outputYes; synthesized agent speech at 22.05 kHz.
Image input or outputNot supported for this model.
Video input or outputNot supported for this model.
StreamingYes; streaming is central to its real-time voice-agent design.
Tool or function callingNot documented for this exact model.
Structured JSON modeNo separate JSON-mode capability is published for this model.

Although the system produces both text and audio, its defining output is conversational speech rather than images, video, or other non-text media. The model should therefore be evaluated as a specialized voice system, not as a general multimodal assistant.

Main strengths

The clearest strength is low-latency, interruptible interaction. A voice agent built around Nemotron 3 VoiceChat can be designed to respond while a conversation is still unfolding, which is useful for spoken interfaces where waiting for a complete user turn feels unnatural. This is especially relevant to assistants, customer-service prototypes, robotics research, and other applications where timing matters as much as answer quality.

The unified speech-to-speech design can also simplify experimentation. Instead of treating speech recognition, language generation, and speech synthesis as entirely unrelated stages, developers can evaluate a system designed around the complete conversational loop. The model returns both text and audio, which can help applications display transcripts, log interactions, or use text for downstream interface elements while playing the spoken response.

NVIDIA documents the model for NVIDIA GPU-based deployment and specifically optimizes it for Hopper hardware, especially H100 GPUs, with vLLM on Linux. That focus may be valuable for organizations already operating NVIDIA data-center infrastructure and looking for a voice model that fits their existing serving environment.

Limitations and unknown specifications

Nemotron 3 VoiceChat is explicitly an early-access evaluation model. It should not be treated as a finished, broadly supported production service without verifying access conditions, hardware requirements, software compatibility, and support terms with NVIDIA.

NVIDIA does not publish a general context-window limit or maximum output-token limit for this exact model. Those values should remain unknown rather than being inferred from the language backbone or from other Nemotron products. The model also has no published knowledge-cutoff date. Its primary purpose is real-time speech interaction, not document-scale analysis or knowledge-intensive text work.

Multilingual support is not established in the supplied documentation, so applications requiring multiple languages should not assume that the model meets those needs. Image and video understanding are outside its documented scope. Tool use and function calling are also not documented for this exact model, which means an application should not assume that the model can directly call business systems, search services, or external APIs without an additional orchestration layer.

NVIDIA documents a guided fine-tuning path through the early-access program, but this does not necessarily mean that unrestricted self-service fine-tuning is generally available. Prospective users should confirm eligibility and the exact training workflow before designing a deployment around customization.

Speed, cost, and deployment trade-offs

The supplied research rates the model’s speed highly on an editorial scale because streaming and real-time voice interaction are central design goals. That score is an evaluation aid, not a benchmark published by NVIDIA. Actual latency will depend on the GPU, serving configuration, audio buffering, network conditions, concurrency, and the surrounding application.

The model is listed as available through a free NVIDIA NIM trial endpoint for evaluation. NVIDIA has not published a general production price for this exact model in the supplied sources. Therefore, “free” should be understood as trial or evaluation access, not as a confirmed unlimited production offering. Self-hosted deployment can also involve substantial infrastructure costs because the documented target environment is enterprise-grade NVIDIA hardware, particularly H100-class systems.

Its cost profile is consequently different from that of a lightweight hosted speech API. Nemotron 3 VoiceChat may be attractive to an organization that already owns or rents NVIDIA infrastructure and needs control over model serving. A managed speech-to-speech service may be more practical for a small team that wants predictable usage pricing and does not need to operate GPU infrastructure.

Reasoning and coding suitability

This model can generate language as part of a spoken conversation, but its design priority is responsive dialogue rather than extended reasoning. An editorial reasoning score of 5 out of 10 indicates a middle-of-the-range assessment for general reasoning, not a provider-published benchmark or guarantee. The model should not be selected primarily for complex research, long multi-step analysis, or large document synthesis.

Similarly, an editorial coding score of 3 out of 10 reflects its limited suitability for programming-focused work. Nemotron 3 VoiceChat may be able to discuss code in a voice interaction, but the supplied research does not support treating it as a coding specialist. A text-focused coding model or a general-purpose language model with documented tool integration would usually be a better fit for software development workflows.

Best use cases

  • Real-time voice agents that need natural interruption handling.
  • Speech-to-speech research and evaluation.
  • Conversational interfaces deployed on NVIDIA GPU infrastructure.
  • Enterprise prototypes where audio, transcription, and generated text are all useful outputs.
  • Interactive assistants that need streaming responses rather than delayed turn-based answers.

For example, a voice-agent prototype could stream a user’s 16 kHz audio to the model, display the returned transcription, play the generated 22.05 kHz speech, and use the agent text for captions or application logs. The exact interruption and turn-management behavior will still depend on the client and serving implementation, so developers should validate it in the target environment.

When to choose Nemotron 3 VoiceChat

Choose Nemotron 3 VoiceChat when the central requirement is a responsive, interruptible voice conversation and the project can work within NVIDIA’s early-access and hardware expectations. It is particularly relevant when an organization already uses NVIDIA GPUs, wants to evaluate a unified speech-to-speech architecture, or needs text and audio outputs from the same interaction.

Choose another option when the primary workload is text generation, coding, image understanding, long-context analysis, multilingual speech, or direct tool execution. A modular speech pipeline can also be a better choice when the project needs to swap speech-recognition or text-to-speech providers independently. For cost-sensitive deployments, a managed endpoint with transparent production pricing may be preferable to operating a model optimized for H100-class infrastructure.

Bottom line

NVIDIA Nemotron 3 VoiceChat is a specialized early-access model for full-duplex, real-time speech interaction. Its important differentiator is not broad general-purpose intelligence but the combination of streaming speech input, conversational generation, transcription, and synthesized speech in a single voice-agent-oriented system. The main benefits are responsiveness and NVIDIA-oriented deployment; the main uncertainties are early-access availability, infrastructure requirements, and the absence of published context, output-token, production-pricing, and tool-use specifications.


Answers to Frequently Asked Questions

How is Nemotron 3 VoiceChat deployed and accessed?
NVIDIA provides the model through an early-access program and an NVIDIA NIM experience under the serving identity nemotron-voicechat. It is documented for NVIDIA GPU-based deployment, particularly Hopper hardware with vLLM on Linux, and is also listed as available through a free NIM trial endpoint for evaluation.
What are the main limitations of Nemotron 3 VoiceChat?
The model is an early-access evaluation release, and NVIDIA has not published a general context-window limit, maximum output-token limit, knowledge-cutoff date, production price, or documented tool-calling support for this exact model. Multilingual support is also not established, and deployment may require enterprise-grade NVIDIA hardware such as H100 GPUs.
What are the best use cases for NVIDIA Nemotron 3 VoiceChat?
Nemotron 3 VoiceChat is suited to real-time voice agents, speech-to-speech research, conversational interfaces, enterprise prototypes, and interactive assistants that require streaming responses and natural interruption handling, especially on NVIDIA GPU infrastructure.
What is NVIDIA Nemotron 3 VoiceChat?
NVIDIA Nemotron 3 VoiceChat is a 12-billion-parameter, full-duplex speech-to-speech model designed for real-time voice agents. It supports simultaneous two-way audio interaction, allowing users to interrupt the model instead of waiting for a turn-based response to finish.
What input and output formats does Nemotron 3 VoiceChat support?
The model accepts text prompts and streaming user audio sampled at 16 kHz. It produces agent-generated text, a transcription of the user’s speech, and synthesized agent speech sampled at 22.05 kHz. Image and video inputs or outputs are not supported.


Sources 5
Provider

About NVIDIA AI