Model catalog

NVIDIA AI Models

Browse the AI models associated with NVIDIA AI. Compare current and historical models by family, capabilities, context window, availability and intended use.

47 models tracked
47 Total models
39 Model families
4 Model types
47 Current / accessible
All models

NVIDIA AI model catalog

NVIDIA AI logo
Active Speaker Detection

Active Speaker Detection

Real-time or batch speaker identification and active-speaker tagging in broadcast, video communication, dubbing, localization, conferencing, and media analytics workflows.

Type Multimodal
Reasoning 1/10
Speed 8/10
Multimodal Audio input Video input
Status

Current; downloadable NVIDIA NIM endpoint

View model →
NVIDIA AI logo
AI4M Relighting

Relighting

Real-time video relighting, virtual production, media effects, HDR-based lighting changes, and foreground/background compositing.

Type Other
Reasoning 1/10
Speed 8/10
Multimodal Image input Video input
Status

Current and accessible through NVIDIA AI for Media and NVIDIA NIM documentation; version 1.1.0 is documented.

View model →

Autonomous-driving research, trajectory prediction, interpretable motion planning, navigation-conditioned driving, visual question answering and safety-oriented model evaluation

Type Reasoning
Reasoning 8/10
Speed 5/10
Multimodal Image input Video input
Status

Current; open weights available; also available through NVIDIA Alpamayo 1.5 NIM

View model →

Multilingual offline speech transcription, speech translation, and NeMo-based ASR research or deployment

Type Other
Reasoning 1/10
Speed 6/10
Audio input
Status

Active; open-weight checkpoint and NVIDIA deployment options available

View model →

Fast English speech-to-text transcription, NeMo experimentation, domain fine-tuning, and Riva-based ASR deployment

Type Other
Reasoning 1/10
Speed 9/10
Audio input Fine-tuning Streaming
Status

Available downloadable checkpoint; older NeMo ASR model

View model →
NVIDIA AI logo
Cosmos-Transfer2.5

Cosmos-Transfer2.5-2B

Controllable video world generation, robotics sim-to-real augmentation, autonomous-vehicle simulation, and Physical AI synthetic-data generation

Type Multimodal
Reasoning 1/10
Speed 3/10
Multimodal Image input Video input
Status

Available; legacy relative to Cosmos 3; original repository under limited maintenance

Input No official token-based hosted price; NVIDIA lists a free NIM endpoint. Downloadable model is intended for self-hosted deployment.
Output No official token-based hosted price; video-generation costs depend on deployment infrastructure and hardware.
View model →

Edge physical AI, robotics, visual reasoning, world simulation, video generation, and action-policy prototyping

Type Multimodal
Context 131K
Reasoning 7/10
Speed 9/10
Multimodal Image input Video input
Status

Current; open model; gated Hugging Face access

View model →

Physical AI, robotics, autonomous-vehicle simulation, multimodal world generation, future-state prediction, action reasoning, and synthetic training data

Type Multimodal
Reasoning 7/10
Speed 6/10
Multimodal Image input Audio input
Status

Current; downloadable open-weight model and available through an NVIDIA NIM endpoint

View model →

High-quality Physical AI simulation, synthetic-data generation, robotics and autonomous-vehicle research, multimodal world modeling, and teacher-model distillation

Type Multimodal
Context 262K
Reasoning 8/10
Speed 2/10
Multimodal Image input Audio input
Status

Current; open-weight model; commercially and non-commercially usable

View model →

Physical-world video and image understanding, robotic perception, embodied-agent planning, spatial-temporal reasoning, and Physical AI research

Type Reasoning
Context 256K
Reasoning 8/10
Speed 6/10
Multimodal Image input Video input
Status

Current; downloadable and hosted NIM endpoint

Input Free downloadable endpoint; hosted pricing not specified in the reviewed NVIDIA sources
Output Free downloadable endpoint; hosted pricing not specified in the reviewed NVIDIA sources
View model →
NVIDIA AI logo
FourCastNet

FourCastNet

Rapid global weather forecasting, ensemble simulation, climate-risk analysis, renewable-energy forecasting, and scientific weather-model research

Type Other
Reasoning 0/10
Speed 9/10
Status

Current and downloadable; available through NVIDIA Earth-2 FourCastNet NIM and related deployment tooling

Input Not publicly listed as a token-priced model; downloadable model and NIM access are subject to NVIDIA access and licensing terms
Output Not publicly listed as a token-priced model; deployment and hosted inference pricing may depend on the selected NVIDIA service or infrastructure
View model →
NVIDIA AI logo
GenMol

GenMol

De novo molecular design, fragment-constrained generation, linker design, scaffold decoration, hit generation, and lead optimization

Type Other
Context 512
Reasoning 2/10
Speed 7/10
Status

Current; downloadable model and available as an NVIDIA NIM

Input No public per-token model price found; downloadable weights are available under the NVIDIA Open Model License. NIM access is subject to NVIDIA API or deployment terms.
Output No public per-token model price found; NIM output pricing is not specified in the reviewed official documentation.
View model →

Humanoid robot manipulation, cross-embodiment policy learning, robot demonstration fine-tuning, physical AI research, and action-sequence deployment

Type Multimodal
Reasoning 4/10
Speed 6/10
Multimodal Image input Video input
Status

Current; general availability

View model →

Quantum-computing calibration plot interpretation, QPU bring-up and retuning workflows, experiment diagnosis, parameter extraction, fit-quality assessment, and domain-specific calibration agents.

Type Multimodal
Context 262K
Reasoning 7/10
Speed 6/10
Multimodal Image input Fine-tuning
Status

Current; available through NVIDIA Build, NVIDIA NIM, and downloadable checkpoints

View model →
NVIDIA AI logo
LipSync

LipSync

Generative lip dubbing, multilingual video localization, broadcasting, conferencing, and digital-human facial animation

Type Multimodal
Reasoning 1/10
Speed 8/10
Multimodal Image input Audio input
Status

Current; downloadable model and NVIDIA LipSync NIM; private access may be required for some workflows

View model →
NVIDIA AI logo
Llama Nemotron Embed VL

Llama Nemotron Embed VL 1B v2

Multimodal semantic search, visual document retrieval, question-answer retrieval, vector databases, and retrieval-augmented generation

Type Embedding
Context 10K
Reasoning 2/10
Speed 8/10
Multimodal Image input
Status

Current and downloadable

View model →
NVIDIA AI logo
Llama Nemotron Rerank VL

Llama Nemotron Rerank VL 1B v2

Reranking text, document images, and image-text candidates in visual search, multimodal RAG, and question-answering retrieval pipelines

Type Multimodal
Context 8K
Reasoning 2/10
Speed 7/10
Multimodal Image input
Status

Current; downloadable and available through NVIDIA NIM and retrieval APIs

View model →

Multilingual voice agents, accessibility, narration, audiobooks, dubbing, localization and interactive speech applications

Type Other
Reasoning 1/10
Speed 8/10
Media output Fine-tuning Streaming
Status

Current

View model →

Low-latency multilingual text translation, speech-translation pipelines, and self-hosted NVIDIA GPU deployments

Type Other
Reasoning 1/10
Speed 8/10
Status

Current and downloadable; available through NVIDIA NIM and NVIDIA Riva

View model →
NVIDIA AI logo
MolMIM

MolMIM

Small-molecule generation, molecular embeddings, chemical-space exploration, lead optimization, and oracle-guided drug-design workflows

Type Other
Context 128
Status

Available through NVIDIA BioNeMo Framework and NVIDIA NIM; research and development model

View model →

Detecting table cells, rows, columns, and merged-cell structure in document images for OCR alignment, table reconstruction, document ingestion, and retrieval systems

Type Other
Reasoning 1/10
Speed 7/10
Image input
Status

Current; downloadable and available through NVIDIA NIM/Build

View model →

Agentic reasoning, long-context analysis, coding, tool-calling workflows, retrieval-augmented generation, collaborative agents, and high-volume enterprise inference.

Type Reasoning
Context 1M
Reasoning 9/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Current; open-weight; available through NVIDIA NIM, NVIDIA's hosted API trial endpoint, downloadable checkpoints, and self-hosted deployments.

View model →

Frontier reasoning, complex coding, long-context analysis, enterprise RAG, tool-using agents, and multi-agent workflows

Type Reasoning
Context 1M
Reasoning 10/10
Speed 7/10
Tool use Fine-tuning Streaming
Status

Current; open-weight model with BF16 and NVFP4 checkpoints

View model →

Multilingual semantic search, dense retrieval, RAG, agentic retrieval, code search, and vector-based document matching

Type Embedding
Context 33K
Fine-tuning
Status

Current; available through NVIDIA NIM and downloadable Hugging Face weights

View model →

Multimodal document intelligence, OCR, long-video and audio understanding, cross-modal reasoning, voice agents, and self-hosted enterprise inference

Type Multimodal
Context 262K
Reasoning 8/10
Speed 8/10
Multimodal Image input Audio input
Status

Current; generally available open-weight checkpoint with hosted NVIDIA NIM access

View model →
NVIDIA AI logo
Nemotron 3 VoiceChat

NVIDIA Nemotron 3 VoiceChat

Real-time full-duplex voice agents, interruptible conversational interfaces, speech-to-speech research, and NVIDIA GPU-based enterprise voice applications

Type Multimodal
Reasoning 5/10
Speed 9/10
Multimodal Audio input Media output
Status

Early access; available for evaluation through NVIDIA NIM and qualified access programs

Input Free endpoint for NVIDIA NIM trial access; no general production price published
Output Free endpoint for NVIDIA NIM trial access; no general production price published
View model →

High-volume agent execution, long-running AI agents, reasoning, coding assistance, RAG, chat, low-latency text generation, and domain customization

Type Reasoning
Context 1M
Reasoning 8/10
Speed 9/10
Tool use Fine-tuning Streaming
Status

Current open-weight model; NIM container available for early-access evaluation

View model →

Low-latency live speech-to-text, voice interfaces, live captions, and continuous audio streams

Type Other
Reasoning 1/10
Speed 9/10
Multimodal Audio input Streaming
Status

Current and accessible through NVIDIA Speech NIM; streaming-only deployment

View model →
NVIDIA AI logo
Nemotron Content Safety

Nemotron 3.5 Content Safety

Multilingual text-and-image moderation, LLM and VLM guardrails, response safety evaluation, and custom enterprise safety policies

Type Other
Context 128K
Reasoning 5/10
Speed 8/10
Multimodal Image input Streaming
Status

Current; open-weight model and downloadable NVIDIA NIM; latest NGC container version bf16-v1.1 as of September 17, 2026

View model →
NVIDIA AI logo
Nemotron Graphic Elements

nemotron-graphic-elements-v1

Detecting and localizing chart titles, axis labels, legends, mark labels, and value labels in document images

Type Other
Reasoning 1/10
Speed 7/10
Image input
Status

Current and downloadable

View model →
NVIDIA AI logo
Nemotron OCR

Nemotron OCR v1

English OCR, document ingestion, layout-aware text extraction, multimodal retrieval, RAG preprocessing, and enterprise document intelligence

Type Other
Reasoning 1/10
Speed 8/10
Image input
Status

Available; English-only OCR model; newer Nemotron OCR v2 is available for updated English and multilingual OCR deployments

View model →

Multilingual OCR, scanned documents, forms, reports, charts, tables, image-based search, document ingestion, and retrieval-augmented generation preprocessing

Type Other
Reasoning 1/10
Speed 9/10
Image input
Status

Current; downloadable model and available through NVIDIA NIM and NVIDIA-hosted services

View model →
NVIDIA AI logo
Nemotron Page Elements

nemotron-page-elements-v3

Document page-layout detection before OCR, table extraction, indexing, enterprise document processing, and multimodal RAG pipelines

Type Other
Reasoning 1/10
Speed 7/10
Image input
Status

Current; downloadable and available through NVIDIA NIM

View model →

Document OCR, layout analysis, table and chart extraction, retrieval pipelines, document indexing, and multimodal data curation

Type Multimodal
Reasoning 2/10
Speed 7/10
Multimodal Image input
Status

current

View model →
NVIDIA AI logo
NVIDIA Ising Calibration

NVIDIA-Ising-Calibration-1-35B-A3B

Quantum-computing calibration plot analysis, experiment interpretation, fit-quality assessment, parameter extraction, and technical research workflows

Type Multimodal
Context 262K
Reasoning 7/10
Speed 6/10
Multimodal Image input Streaming
Status

Current; downloadable open-weight model

View model →
NVIDIA AI logo
NVIDIA Maxine Eye Contact

Eye Contact

Gaze correction in video conferencing, telepresence, digital-human applications, and video-processing pipelines

Type Other
Reasoning 1/10
Speed 8/10
Multimodal Image input Video input
Status

Current; downloadable NVIDIA NIM model

View model →

English speech-to-text transcription, local ASR, offline processing, RAG audio extraction, and customized NeMo deployments

Type Other
Reasoning 1/10
Speed 8/10
Audio input Fine-tuning Streaming
Status

Available open-weight checkpoint

View model →
NVIDIA AI logo
Parakeet CTC

Parakeet CTC 0.6B

Local English speech transcription, offline ASR, long-form audio processing, and custom ASR fine-tuning

Type Other
Reasoning 1/10
Speed 9/10
Audio input Fine-tuning
Status

Current and accessible open-weight model

View model →

Taiwanese Mandarin speech-to-text, Mandarin-English code-switching, live captions, voice interfaces, and NVIDIA GPU-hosted streaming or offline transcription.

Type Other
Reasoning 1/10
Speed 8/10
Audio input Streaming
Status

Listed in current NVIDIA Speech NIM documentation and model catalog; the associated NGC container is marked no longer supported.

View model →

Mandarin-English automatic speech recognition, code-switched transcription, streaming transcription, and self-hosted enterprise speech-to-text

Type Other
Reasoning 1/10
Speed 8/10
Audio input Streaming
Status

Current and downloadable; available through NVIDIA NIM and Riva interfaces

Input No public per-token price; downloadable model and NVIDIA service terms apply
Output No public per-token price; downloadable model and NVIDIA service terms apply
View model →

High-speed English speech transcription, captions, meeting transcription, audio search, and timestamped media workflows

Type Other
Reasoning 1/10
Speed 10/10
Audio input Fine-tuning Streaming
Status

Current open-weight model; publicly available through Hugging Face and NVIDIA NeMo

View model →

Multilingual sentence and document translation, localization, marketing content, and developer translation workflows

Type Other
Context 8K
Reasoning 2/10
Speed 7/10
Status

Current and accessible through NVIDIA NIM and downloadable model weights

Input Free hosted endpoint listed on Build.NVIDIA.com; no self-hosted per-token price specified
Output Free hosted endpoint listed on Build.NVIDIA.com; no self-hosted per-token price specified
View model →

Local or self-hosted multilingual sentence and document translation across English and 36 non-English languages

Type Other
Context 8K
Reasoning 3/10
Speed 8/10
Status

Current and publicly downloadable; supported by NVIDIA NIM

View model →
NVIDIA AI logo
StreamPETR

StreamPETR

Camera-only multi-view 3D perception, autonomous-driving scene analysis, bird's-eye-view visualization, and object tracking

Type Other
Speed 8/10
Multimodal Image input Video input
Status

Current NVIDIA NIM endpoint; free endpoint access requires an NVIDIA API key

Input Free endpoint
View model →
NVIDIA AI logo
Studio Voice

Studio Voice

Real-time enhancement of speech captured with low-quality microphones in noisy or reverberant environments, including broadcast, conferencing, telecommunications, and media production.

Type Other
Reasoning 1/10
Speed 9/10
Audio input Streaming
Status

Current and available through NVIDIA NIM, hosted preview services, and NVIDIA audio software; downloadable deployment may require applicable NVIDIA licensing or subscription access.

View model →
NVIDIA AI logo
Synthetic Video Detector

NVIDIA Synthetic Video Detector

AI-generated video detection, media authentication, digital forensics, content moderation, and media-integrity monitoring

Type Other
Reasoning 1/10
Speed 8/10
Video input Streaming
Status

Current; downloadable NIM and managed trial endpoint, with self-hosted deployment requiring AI for Media Private Access

Input Free managed trial endpoint; self-hosted and enterprise licensing terms may apply
View model →
NVIDIA AI logo
Video Super Resolution

Video Super Resolution NIM

Professional video upscaling, broadcast enhancement, streaming pipelines, pre-encoding optimization, denoising, and deblurring

Type Other
Reasoning 1/10
Speed 8/10
Multimodal Video input Media output
Status

Current; downloadable NVIDIA NIM

View model →