NVIDIA Maxine Eye Contact

Eye Contact

by NVIDIA AI · Current; downloadable NVIDIA NIM model

NVIDIA Eye Contact is a specialized Maxine model for estimating gaze direction and synthesizing redirected eye regions that appear to look toward the camera. It supports downloadable NIM deployment and service-level MP4/H.264 processing for video conferencing, telepresence, recorded presentations, and digital-human workflows.

Image generation Reasoning Coding
NVIDIA Eye Contact uses an encoder-decoder neural network to analyze an eye-region image, estimate gaze direction, and synthesize a redirected eye patch. The model can be deployed through NVIDIA NIM on supported NVIDIA GPUs and can also be used in a video-processing service that accepts MP4/H.264 input.
Outputs

What Eye Contact can produce

Image generation
Inputs

What it can understand

Images Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family NVIDIA Maxine Eye Contact
Model type Other
Release date 2023-01-21
Status Current; downloadable NVIDIA NIM model
Knowledge cutoff notes

NVIDIA does not publish a conventional textual knowledge cutoff for this vision-processing model. It is trained for gaze estimation and redirection rather than factual knowledge retrieval.

Model notes

The canonical model identity is Eye Contact, with model ID maxine-eye-contact. NVIDIA's model card describes direct neural-network inputs as a normalized 256 × 64 RGB eye patch and a two-element pitch/yaw angle vector. Direct outputs are a redirected RGB eye patch, an estimated angle vector, and eye landmarks. NVIDIA's NIM API demonstrates MP4/H.264 video processing and can produce a redirected video at the service level. The model is distributed under NVIDIA Maxine evaluation and deployment terms; no public token-based model pricing was identified.

Model guide

NVIDIA Eye Contact: AI Gaze Redirection for Video Conferencing

NVIDIA Eye Contact is a specialized NVIDIA Maxine model and downloadable NIM microservice that estimates gaze direction and redirects a person's eyes toward the camera. It is designed for video conferencing, telepresence, digital humans, and related video-processing workflows rather than general-purpose conversation or content generation.

What is NVIDIA Eye Contact?

NVIDIA Eye Contact is a specialized computer-vision model in the NVIDIA Maxine family. Its purpose is narrow and practical: it changes the apparent direction of a person's gaze so that the eyes look more toward the camera. This can make remote presentations, video calls, telepresence systems, and digital-human applications appear more natural when a participant is looking at a screen instead of directly into the camera.

The model is not a chatbot, language model, or general video-understanding system. It does not generate written answers, hold a conversation, write code, search the web, or interpret unrestricted video content. Instead, it performs gaze estimation and eye-gaze redirection on a defined eye-region input.

NVIDIA makes Eye Contact available as a downloadable NVIDIA NIM microservice, with documentation and a demonstration through NVIDIA Build. NIM is NVIDIA's packaging and serving approach for deploying AI inference services. In this case, the service can be hosted or self-managed on compatible NVIDIA hardware.

What the model does

At a high level, Eye Contact receives an image of the eye region and a requested redirection angle. It estimates the existing gaze, modifies the relevant internal representation, and generates a new eye patch in which the gaze has been redirected. The result is intended to be composited back into the surrounding face and video frame by the application using it.

The model uses an encoder-decoder architecture. The encoder converts the eye image into latent representations, meaning compact internal features that describe information such as the subject, head pose, environmental conditions, and gaze direction. A transformation is applied to the gaze-related information, and the decoder produces the redirected eye region.

This design makes the model useful for a focused image transformation, but it also defines its boundaries. An application must generally manage the surrounding video pipeline, including face and eye detection, frame extraction, eye-patch preparation, compositing, video decoding, and output encoding.

Inputs and outputs

The documented neural-network interface uses a normalized RGB eye patch measuring 256 × 64 pixels. The input also includes a two-element angle vector containing pitch and yaw values in radians. Pitch represents vertical gaze movement, while yaw represents horizontal movement.

The direct model outputs are:

  • A redirected RGB eye image.
  • An estimated gaze-angle vector.
  • Eye landmarks represented as fiducial points.

At the NIM service level, NVIDIA documents MP4 video encoded with H.264 as an input format. The service can return processed video, but this should not be confused with the neural network's direct output: the underlying model produces an eye patch and related gaze and landmark data, while the service can wrap that operation in a broader video workflow.

The supplied specifications do not identify a conventional context window, maximum token count, maximum generated-text length, or other language-model output limit. Those limits are not applicable to Eye Contact's direct image-processing interface.

Where it fits in NVIDIA's catalog

Eye Contact belongs to NVIDIA Maxine, a family of technologies for audio, video, and communication-enhancement workloads. Within that family, Eye Contact occupies a specific gaze-correction role rather than serving as a general-purpose generative model.

Its current documented distribution is through NVIDIA NIM and an associated downloadable container. The support documentation lists NVIDIA GPU architectures including Volta, Turing, Ampere, Ada, and Blackwell. Supported configurations include professional and consumer GPUs with Tensor Cores, and NVIDIA recommends NVENC and NVDEC support for video workloads.

This hardware orientation is important when evaluating the model. Eye Contact is not primarily a consumer application that hides infrastructure requirements behind a simple account. A deployment normally requires compatible NVIDIA hardware, the appropriate software environment, and an application capable of preparing and processing the video stream.

Strengths and practical benefits

  • Focused output: The model is purpose-built for gaze correction, so its output corresponds directly to a common video-conferencing problem.
  • Useful technical outputs: In addition to the redirected eye patch, it provides estimated gaze angles and eye landmarks that can support a larger processing pipeline.
  • Video-service integration: The NIM interface supports MP4/H.264 video processing, reducing the amount of service infrastructure an application must build around the model.
  • Deployable on NVIDIA hardware: The documented support across several NVIDIA GPU generations provides options for workstation, professional, and data-center deployment, subject to the relevant compatibility requirements.
  • Suitable for real-time-oriented workflows: NVIDIA positions the model for video conferencing and related communication applications. Actual performance will depend on the GPU, video resolution, pipeline design, and deployment configuration; the supplied research does not provide a universal latency benchmark.

Its main advantage is specialization. An organization needing a gaze-redirection component does not need to use a large vision-language model for a task that can be expressed as a constrained image transformation.

Limitations and trade-offs

Eye Contact is not a general-purpose AI assistant. It has no documented text-generation output, conversational reasoning, web search, function calling, coding capability, speech processing, image generation, or general video-understanding capability. The model's reasoning and coding scores in the supplied catalog are editorial database fields, not provider-published measures, and they should not be interpreted as evidence that the model performs those tasks.

The input format is also constrained. The documented model expects a correctly prepared eye patch and a pitch/yaw instruction rather than an arbitrary image or full video frame. A production application must locate the face and eyes, crop the appropriate region, normalize the image, apply the model output to the original frame, and handle failures such as poor lighting, occlusion, unusual poses, or tracking errors. The supplied documentation does not establish a universal quality threshold for those conditions.

The model's hardware requirements can make it less convenient than a cloud video API or a software solution designed for broader device compatibility. Conversely, a self-managed NVIDIA deployment may be preferable where an organization needs control over processing location, pipeline integration, or recurring service usage. No public token-based or per-minute price was identified for the model, so total cost depends on hardware, hosting, licensing, and implementation choices.

Pricing and availability

No public model price, token price, or recurring subscription price was identified in the supplied NVIDIA sources. Eye Contact is described as a downloadable NIM model or container rather than as a conventional consumer subscription with a simple monthly tier.

NVIDIA's model and demonstration materials apply license terms that govern evaluation and deployment. Organizations should review the applicable NVIDIA Maxine, NIM, container, and hardware requirements before using the model commercially. The absence of a listed API price should not be treated as evidence that production deployment is cost-free.

Supported modalities and capabilities

CapabilitySupported or documented behavior
Image inputYes; normalized RGB eye patch, documented at 256 × 64 pixels
Video inputYes at the NIM service level; MP4/H.264 is documented
Direct image outputYes; redirected RGB eye patch
Video outputAvailable at the service level when processing video
Text input or outputNot documented for the model
Audio input or outputNot documented for the model
ReasoningNot a general reasoning model
CodingNot supported as a model task
Tools or function callingNot documented
Structured text or JSON generationNot documented as a model capability

The model's multimodal nature should therefore be understood as visual input and visual output, not as broad multimodal conversation. It can work with video through the surrounding NIM service, but it does not analyze video in the way a general vision-language model would.

When to choose NVIDIA Eye Contact

Choose Eye Contact when the primary requirement is gaze correction and the deployment environment can support NVIDIA GPUs. It is a good fit for:

  • Video-conferencing software that wants to make screen-based eye contact appear more natural.
  • Recorded presentations where the speaker's gaze should be redirected toward the viewer.
  • Telepresence and digital-human systems that need a dedicated eye-contact component.
  • Media-processing pipelines that already use NVIDIA hardware and need an on-premises or self-managed inference service.

A specialized model is often more practical than a large generative model when the task is limited to one controlled transformation. It can reduce unnecessary model capability and keep the processing objective clear. However, the supplied research does not provide enough benchmark data to claim that it is faster or cheaper than every alternative. Those comparisons depend on hardware, resolution, frame rate, batching, and the amount of surrounding video processing.

When another option may be more appropriate

Use a general vision-language model when the application must describe video, answer questions about scenes, extract information from varied footage, or combine visual input with natural-language reasoning. Use a speech or avatar platform when the main requirement is voice interaction, lip synchronization, or full digital-human generation rather than gaze correction alone.

A cloud-based video-processing service may be easier for teams that do not want to operate NVIDIA infrastructure. A broader facial-animation or media-effects solution may also be preferable when the application needs many transformations beyond the eyes. These alternatives may trade away the focused deployment model of Eye Contact for simpler integration or wider functionality.

Technical summary

PropertyVerified information
ProviderNVIDIA
Model familyNVIDIA Maxine Eye Contact
Model IDmaxine-eye-contact
Primary taskGaze estimation and eye-gaze redirection
ArchitectureConvolutional encoder-decoder network
Direct inputNormalized RGB eye patch and pitch/yaw angle vector
Documented eye-patch size256 × 64 pixels
Direct outputRedirected RGB eye patch, estimated gaze angle, and eye landmarks
Service video formatMP4 encoded with H.264
Inference technologiesTensorRT and Triton
DeploymentDownloadable NVIDIA NIM microservice/container
Public model pricingNot identified

NVIDIA Eye Contact is best understood as an infrastructure-ready component for a specific visual problem. Its value comes from producing a controlled gaze transformation and associated tracking data, not from broad generative or conversational intelligence. Teams evaluating it should start with the quality of their eye-detection and compositing pipeline, compatible GPU availability, licensing requirements, and the cost of operating the surrounding video system.


Answers to Frequently Asked Questions

What hardware and deployment options does NVIDIA Eye Contact require?
NVIDIA Eye Contact is distributed as a downloadable NVIDIA NIM microservice or container and is intended for compatible NVIDIA hardware. Documented GPU architectures include Volta, Turing, Ampere, Ada, and Blackwell, with Tensor Cores and recommended NVENC and NVDEC support for video workloads. A production application must also handle face and eye detection, eye-patch preparation, compositing, video decoding, and output encoding.
Is NVIDIA Eye Contact a general-purpose AI or language model?
No. NVIDIA Eye Contact is a specialized visual transformation model, not a chatbot, language model, or general video-understanding system. It does not generate text, answer questions, search the web, write code, process speech, or provide broad conversational reasoning.
What is NVIDIA Eye Contact used for?
NVIDIA Eye Contact is a computer-vision model for redirecting a person's apparent gaze toward the camera. It is designed for video conferencing, recorded presentations, telepresence systems, and digital-human applications where natural-looking eye contact is important.
What inputs and outputs does NVIDIA Eye Contact support?
The direct model input is a normalized RGB eye patch measuring 256 × 64 pixels, along with a two-element pitch and yaw angle vector in radians. It outputs a redirected RGB eye patch, an estimated gaze-angle vector, and eye landmarks. At the NVIDIA NIM service level, MP4 video encoded with H.264 is also documented.


Sources 6
Provider

About NVIDIA AI