Active Speaker Detection
Real-time or batch speaker identification and active-speaker tagging in broadcast, video communication, dubbing, localization, conferencing, and media analytics workflows.
Current; downloadable NVIDIA NIM endpoint
Browse the AI models associated with NVIDIA AI. Compare current and historical models by family, capabilities, context window, availability and intended use.
Real-time or batch speaker identification and active-speaker tagging in broadcast, video communication, dubbing, localization, conferencing, and media analytics workflows.
Current; downloadable NVIDIA NIM endpoint
Real-time video relighting, virtual production, media effects, HDR-based lighting changes, and foreground/background compositing.
Current and accessible through NVIDIA AI for Media and NVIDIA NIM documentation; version 1.1.0 is documented.
Autonomous-driving research, trajectory prediction, interpretable motion planning, navigation-conditioned driving, visual question answering and safety-oriented model evaluation
Current; open weights available; also available through NVIDIA Alpamayo 1.5 NIM
Multilingual offline speech transcription, speech translation, and NeMo-based ASR research or deployment
Active; open-weight checkpoint and NVIDIA deployment options available
Fast English speech-to-text transcription, NeMo experimentation, domain fine-tuning, and Riva-based ASR deployment
Available downloadable checkpoint; older NeMo ASR model
Controllable video world generation, robotics sim-to-real augmentation, autonomous-vehicle simulation, and Physical AI synthetic-data generation
Available; legacy relative to Cosmos 3; original repository under limited maintenance
Edge physical AI, robotics, visual reasoning, world simulation, video generation, and action-policy prototyping
Current; open model; gated Hugging Face access
Physical AI, robotics, autonomous-vehicle simulation, multimodal world generation, future-state prediction, action reasoning, and synthetic training data
Current; downloadable open-weight model and available through an NVIDIA NIM endpoint
High-quality Physical AI simulation, synthetic-data generation, robotics and autonomous-vehicle research, multimodal world modeling, and teacher-model distillation
Current; open-weight model; commercially and non-commercially usable
Physical-world video and image understanding, robotic perception, embodied-agent planning, spatial-temporal reasoning, and Physical AI research
Current; downloadable and hosted NIM endpoint
Rapid global weather forecasting, ensemble simulation, climate-risk analysis, renewable-energy forecasting, and scientific weather-model research
Current and downloadable; available through NVIDIA Earth-2 FourCastNet NIM and related deployment tooling
De novo molecular design, fragment-constrained generation, linker design, scaffold decoration, hit generation, and lead optimization
Current; downloadable model and available as an NVIDIA NIM
Humanoid robot manipulation, cross-embodiment policy learning, robot demonstration fine-tuning, physical AI research, and action-sequence deployment
Current; general availability
Quantum-computing calibration plot interpretation, QPU bring-up and retuning workflows, experiment diagnosis, parameter extraction, fit-quality assessment, and domain-specific calibration agents.
Current; available through NVIDIA Build, NVIDIA NIM, and downloadable checkpoints
Generative lip dubbing, multilingual video localization, broadcasting, conferencing, and digital-human facial animation
Current; downloadable model and NVIDIA LipSync NIM; private access may be required for some workflows
Multimodal semantic search, visual document retrieval, question-answer retrieval, vector databases, and retrieval-augmented generation
Current and downloadable
Reranking text, document images, and image-text candidates in visual search, multimodal RAG, and question-answering retrieval pipelines
Current; downloadable and available through NVIDIA NIM and retrieval APIs
Multilingual voice agents, accessibility, narration, audiobooks, dubbing, localization and interactive speech applications
Current
Low-latency multilingual text translation, speech-translation pipelines, and self-hosted NVIDIA GPU deployments
Current and downloadable; available through NVIDIA NIM and NVIDIA Riva
Small-molecule generation, molecular embeddings, chemical-space exploration, lead optimization, and oracle-guided drug-design workflows
Available through NVIDIA BioNeMo Framework and NVIDIA NIM; research and development model
Detecting table cells, rows, columns, and merged-cell structure in document images for OCR alignment, table reconstruction, document ingestion, and retrieval systems
Current; downloadable and available through NVIDIA NIM/Build
Agentic reasoning, long-context analysis, coding, tool-calling workflows, retrieval-augmented generation, collaborative agents, and high-volume enterprise inference.
Current; open-weight; available through NVIDIA NIM, NVIDIA's hosted API trial endpoint, downloadable checkpoints, and self-hosted deployments.
Frontier reasoning, complex coding, long-context analysis, enterprise RAG, tool-using agents, and multi-agent workflows
Current; open-weight model with BF16 and NVFP4 checkpoints
Multilingual semantic search, dense retrieval, RAG, agentic retrieval, code search, and vector-based document matching
Current; available through NVIDIA NIM and downloadable Hugging Face weights
Multimodal document intelligence, OCR, long-video and audio understanding, cross-modal reasoning, voice agents, and self-hosted enterprise inference
Current; generally available open-weight checkpoint with hosted NVIDIA NIM access
Real-time full-duplex voice agents, interruptible conversational interfaces, speech-to-speech research, and NVIDIA GPU-based enterprise voice applications
Early access; available for evaluation through NVIDIA NIM and qualified access programs
High-volume agent execution, long-running AI agents, reasoning, coding assistance, RAG, chat, low-latency text generation, and domain customization
Current open-weight model; NIM container available for early-access evaluation
Low-latency live speech-to-text, voice interfaces, live captions, and continuous audio streams
Current and accessible through NVIDIA Speech NIM; streaming-only deployment
Multilingual text-and-image moderation, LLM and VLM guardrails, response safety evaluation, and custom enterprise safety policies
Current; open-weight model and downloadable NVIDIA NIM; latest NGC container version bf16-v1.1 as of September 17, 2026
Detecting and localizing chart titles, axis labels, legends, mark labels, and value labels in document images
Current and downloadable
English OCR, document ingestion, layout-aware text extraction, multimodal retrieval, RAG preprocessing, and enterprise document intelligence
Available; English-only OCR model; newer Nemotron OCR v2 is available for updated English and multilingual OCR deployments
Multilingual OCR, scanned documents, forms, reports, charts, tables, image-based search, document ingestion, and retrieval-augmented generation preprocessing
Current; downloadable model and available through NVIDIA NIM and NVIDIA-hosted services
Document page-layout detection before OCR, table extraction, indexing, enterprise document processing, and multimodal RAG pipelines
Current; downloadable and available through NVIDIA NIM
Document OCR, layout analysis, table and chart extraction, retrieval pipelines, document indexing, and multimodal data curation
current
Quantum-computing calibration plot analysis, experiment interpretation, fit-quality assessment, parameter extraction, and technical research workflows
Current; downloadable open-weight model
Gaze correction in video conferencing, telepresence, digital-human applications, and video-processing pipelines
Current; downloadable NVIDIA NIM model
English speech-to-text transcription, local ASR, offline processing, RAG audio extraction, and customized NeMo deployments
Available open-weight checkpoint
Local English speech transcription, offline ASR, long-form audio processing, and custom ASR fine-tuning
Current and accessible open-weight model
Taiwanese Mandarin speech-to-text, Mandarin-English code-switching, live captions, voice interfaces, and NVIDIA GPU-hosted streaming or offline transcription.
Listed in current NVIDIA Speech NIM documentation and model catalog; the associated NGC container is marked no longer supported.
Mandarin-English automatic speech recognition, code-switched transcription, streaming transcription, and self-hosted enterprise speech-to-text
Current and downloadable; available through NVIDIA NIM and Riva interfaces
High-speed English speech transcription, captions, meeting transcription, audio search, and timestamped media workflows
Current open-weight model; publicly available through Hugging Face and NVIDIA NeMo
Multilingual sentence and document translation, localization, marketing content, and developer translation workflows
Current and accessible through NVIDIA NIM and downloadable model weights
Local or self-hosted multilingual sentence and document translation across English and 36 non-English languages
Current and publicly downloadable; supported by NVIDIA NIM
Camera-only multi-view 3D perception, autonomous-driving scene analysis, bird's-eye-view visualization, and object tracking
Current NVIDIA NIM endpoint; free endpoint access requires an NVIDIA API key
Real-time enhancement of speech captured with low-quality microphones in noisy or reverberant environments, including broadcast, conferencing, telecommunications, and media production.
Current and available through NVIDIA NIM, hosted preview services, and NVIDIA audio software; downloadable deployment may require applicable NVIDIA licensing or subscription access.
AI-generated video detection, media authentication, digital forensics, content moderation, and media-integrity monitoring
Current; downloadable NIM and managed trial endpoint, with self-hosted deployment requiring AI for Media Private Access
Professional video upscaling, broadcast enhancement, streaming pipelines, pre-encoding optimization, denoising, and deblurring
Current; downloadable NVIDIA NIM