LipSync

LipSync

by NVIDIA AI · Current; downloadable model and NVIDIA LipSync NIM; private access may be required for some workflows

NVIDIA LipSync is a specialized audiovisual model for synchronizing a human face with target speech. The core model accepts a face image and mono 16 kHz audio, while LipSync NIM processes video, supports streaming or transactional inference, and offers generic, German, Spanish, and French variants. It is designed for dubbing, localization, digital humans, broadcasting, conferencing, and AR applications rather than text generation or general-purpose reasoning.

Image generation Reasoning Coding
NVIDIA LipSync is built for speech-driven facial animation rather than general-purpose text or image generation. It can help localize video into other languages, synchronize digital-human faces with voice tracks, and adapt broadcast or conferencing footage without rebuilding the entire scene. NVIDIA distributes the model through its AR SDK and LipSync NIM deployment stack, with a generic language-agnostic version and fine-tuned variants for German, Spanish, and French.
Outputs

What LipSync can produce

Image generation
Inputs

What it can understand

Images Audio Video Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family LipSync
Model type Multimodal
Release date 2025-10-17
Status Current; downloadable model and NVIDIA LipSync NIM; private access may be required for some workflows
Knowledge cutoff notes

This is a generative audiovisual model rather than a knowledge-grounded language model. NVIDIA does not publish a textual knowledge-cutoff date for the exact model.

Model notes

The canonical model name is LipSync. The downloadable model card describes the core model as accepting one RGB face image and mono 16 kHz PCM speech audio and producing an RGB image. NVIDIA LipSync NIM extends this workflow to video and speech inputs and can stream output for streamable video. The NIM includes a generic language-agnostic model plus fine-tuned German, Spanish, and French variants. The model card lists 6.13*10^8 parameters, CNN and GAN components, an encoder-decoder architecture, up to 4K image input, and NVIDIA Open Model License terms. Hosted or downloadable access may require NVIDIA private-program approval or the relevant SDK.

Model guide

NVIDIA LipSync: Speech-Driven Lip Dubbing for Images and Video

NVIDIA LipSync is a generative lip-dubbing model that changes a person's mouth movements to match target speech while preserving the subject's identity, pose, background, and overall scene. The core model works with a face image and mono speech audio, while NVIDIA LipSync NIM extends the workflow to video and supports streaming or transactional processing.

What is NVIDIA LipSync?

NVIDIA LipSync is a generative lip-dubbing model that synchronizes visible mouth movements with speech audio. Given a face and a target voice track, it generates updated facial imagery in which the lips are timed to the spoken words. The intended result is a localized or re-voiced performance that retains the original person's appearance, head pose, framing, background, and general scene.

The core model is an audiovisual model rather than a conversational language model. It does not take a written prompt and return an explanation, code, or an answer. Instead, it performs a specialized transformation: speech audio guides changes to a human face.

NVIDIA makes LipSync available through its AR SDK and the NVIDIA LipSync NIM deployment stack. NIM is the video-oriented serving layer around the underlying model. The downloadable model card describes frame-level processing with an RGB face image and speech audio, while the NIM combines this processing across video frames and returns synchronized video.

How NVIDIA LipSync works

The documented architecture combines convolutional neural network components with generative adversarial network components in an encoder-decoder design. In practical terms, the system extracts useful information from the face image and speech signal, then generates a revised facial image at the input resolution.

The core model expects one complete human face in an RGB image and mono PCM speech sampled at 16 kHz. It uses the audio timing and speech-related facial cues to generate a new RGB image. For a video, the NIM applies this process across frames while keeping the output synchronized with the supplied audio track.

NVIDIA documents two NIM processing modes. Streaming mode can process streamable video incrementally, which is relevant to interactive or live-oriented workflows. Transactional mode receives the complete input before inference, making it more suitable for file-based processing. The NIM can also accept optional per-frame speaker bounding-box information when a workflow needs to target a particular speaker or coordinate processing in multi-speaker footage.

Inputs, outputs, and supported modalities

ComponentInputsOutput
Core LipSync modelOne RGB image containing one complete human face and mono 16 kHz PCM speech audioAn RGB face image at the input resolution
LipSync NIMVideo and speech audio, with optional per-frame speaker bounding boxesSpeech-synchronized output video with configurable video and audio settings

The practical input modalities are therefore visual media and speech audio. The model does not document text input, text generation, speech recognition, text-to-speech, or prompt-based image generation. Its output is visual: the underlying model produces an image, while the NIM packages the frame-level results as video.

NVIDIA's model card lists support for image input up to 4K. The supplied documentation does not provide a general context-window figure, maximum token count, or textual output limit because LipSync is not a token-based language model. It also does not publish a universal maximum video duration in the supplied research; usable duration will depend on the selected NIM configuration, hardware, stream mode, and deployment environment.

Language variants and deployment

LipSync includes a generic language-agnostic model and language-specific fine-tuned variants for German, Spanish, and French. A deployment selects one model through the NIM model-selector configuration, so one container serves one selected variant at a time. This is important for multilingual localization pipelines: the available variants are not described as one universal model that dynamically switches between all supported languages within a single deployment.

The model is designed for NVIDIA GPU-accelerated systems. The documented hardware coverage includes Turing, Ampere, Lovelace or Ada, Hopper, and Blackwell architectures, with configurations for professional and consumer RTX GPUs. NVIDIA documentation also lists supported Linux distributions and Windows versions. Exact performance and compatibility depend on the GPU, driver, operating system, input resolution, and deployment configuration.

The downloadable model and NIM are part of NVIDIA's broader AI catalog, but LipSync has a narrower role than general-purpose NVIDIA inference services. Its purpose is facial synchronization for audiovisual media, not broad language, retrieval, coding, or agent workflows.

Strengths and practical use cases

LipSync's main strength is focused audiovisual transformation. It is useful when the original performance and scene should remain visually consistent while the speech track changes.

  • Video localization and dubbing: A translated voice track can be paired with mouth movements that better match the new language, reducing the visible mismatch between audio and facial motion.
  • Digital humans: A rendered or recorded face can be synchronized with speech for virtual presenters, avatars, and other character-based applications.
  • Broadcasting: Production teams can explore alternate language or voice versions while preserving the original camera composition.
  • Video conferencing: The streaming NIM mode is relevant to applications that need incremental processing rather than waiting for a complete file.
  • AR SDK workflows: Developers using NVIDIA's augmented-reality tooling can use LipSync as a specialized facial-animation component.

Another practical advantage is that the model targets mouth movement rather than attempting to regenerate the entire scene. That focus can help preserve the surrounding background and the subject's general identity, although real-world output should still be checked for artifacts and synchronization errors.

Limitations and quality considerations

The single-face expectation is a significant limitation. The core model is documented around one complete human face, so footage with several people, partial faces, heavy occlusion, unusual framing, or rapid changes may require additional processing or may not be a good fit. The NIM's optional speaker bounding-box input can help target faces in more complex workflows, but the supplied documentation does not establish unrestricted multi-face editing capability.

LipSync changes visible facial motion; it is not a complete dubbing production system. It does not translate text, generate a replacement voice, recognize speech, or write a script. A localization pipeline still needs separate translation, voice or audio production, timing preparation, and quality review.

Users should also evaluate identity preservation, mouth and teeth artifacts, facial expressions, synchronization quality, language-specific results, and consent. The model card's commercial-use statement does not remove the need to obtain permission from people whose likeness or performance is processed. Misleading or unauthorized alteration of a person's speech and facial movements can create legal, ethical, and reputational risks.

Reasoning, coding, and tool support

LipSync has no general reasoning capability in the language-model sense. It does not analyze a question, plan a multi-step answer, or provide a reasoning trace. It also is not a coding model and does not generate or execute software code as part of its documented function.

The model does not expose general-purpose function calling or tool use. Its operational interface is the media-processing workflow provided by the AR SDK or LipSync NIM. Streaming is supported at the NIM level, but that should not be confused with conversational tool use or an agent API.

Pricing, access, and license

No public per-image, per-minute, per-request, or recurring subscription price is supplied for NVIDIA LipSync. The model is listed as downloadable and governed by the NVIDIA Open Model License, and the model card describes it as ready for commercial use. However, access to a hosted or downloadable workflow may require NVIDIA private-program approval or the relevant NVIDIA SDK and deployment components.

Because LipSync is primarily deployed on NVIDIA GPU infrastructure, the total cost of use may include hardware, cloud GPU time, storage, engineering, and operational support. The supplied research does not establish a fixed cost comparison between streaming and transactional deployments. In general, streaming can reduce waiting for complete-file processing, while transactional processing may be simpler for batch jobs; the actual speed and cost trade-off depends on the selected GPU and workload.

When to choose NVIDIA LipSync

Choose NVIDIA LipSync when the central requirement is to synchronize a visible human face with supplied speech while keeping the surrounding visual performance largely intact. It is a particularly appropriate option for NVIDIA GPU-based teams building multilingual dubbing, digital-human, broadcast, conferencing, or AR applications that need a deployable audiovisual component rather than a general-purpose AI assistant.

Another option may be more appropriate when the task begins with text and requires translation, script generation, speech synthesis, or conversational reasoning. A general video-generation system may also be a better fit when the goal is to create entirely new scenes, characters, or camera movements rather than modify mouth motion in an existing face. LipSync is similarly not the right specialized choice for speech recognition, voice cloning, or unrestricted multi-person video editing.

Its speed and cost profile should be assessed as a GPU deployment decision rather than a simple consumer subscription comparison. NVIDIA documents support for streaming and several GPU generations, but the supplied material does not include benchmark figures or a universal throughput guarantee. Teams should test their target resolution, language variant, video length, and hardware before committing to production.

Bottom line

NVIDIA LipSync is a focused speech-driven facial-animation model. The core model turns a face image and speech audio into a synchronized face image, while LipSync NIM extends the process to video and supports streaming or complete-file inference. Its value comes from preserving an existing subject and scene during dubbing or speech replacement, not from broad generative or conversational capability.


Answers to Frequently Asked Questions

What are the main limitations of NVIDIA LipSync?
The core model is designed for one complete human face and may require additional processing for multi-person footage, partial faces, occlusion, unusual framing, or rapid changes. It does not provide unrestricted multi-face editing, general reasoning, code generation, voice cloning, or scene generation. Users should also evaluate synchronization, identity preservation, facial artifacts, consent, and legal or ethical risks.
Which languages and deployment modes does NVIDIA LipSync support?
LipSync includes a language-agnostic model and language-specific variants for German, Spanish, and French. NVIDIA LipSync NIM supports streaming mode for incremental processing and transactional mode for complete-file inference. One deployment selects one model variant at a time.
Does NVIDIA LipSync translate text or generate speech?
No. NVIDIA LipSync is a facial-synchronization model, not a translation, speech-recognition, text-to-speech, or conversational AI system. A complete dubbing workflow requires separate tools for translation, voice or audio production, timing preparation, and quality review.
What is NVIDIA LipSync used for?
NVIDIA LipSync synchronizes visible mouth movements with supplied speech audio while preserving the subject’s general appearance, pose, framing, background, and scene. Common uses include video dubbing and localization, digital humans, broadcasting, video conferencing, and AR applications.
What inputs and outputs does NVIDIA LipSync support?
The core model accepts one RGB image containing a complete human face and mono PCM speech audio sampled at 16 kHz, then produces a synchronized RGB face image. NVIDIA LipSync NIM accepts video and speech audio, with optional per-frame speaker bounding boxes, and returns synchronized output video.


Sources 5
Provider

About NVIDIA AI