What is NVIDIA Eye Contact?
NVIDIA Eye Contact is a specialized computer-vision model in the NVIDIA Maxine family. Its purpose is narrow and practical: it changes the apparent direction of a person's gaze so that the eyes look more toward the camera. This can make remote presentations, video calls, telepresence systems, and digital-human applications appear more natural when a participant is looking at a screen instead of directly into the camera.
The model is not a chatbot, language model, or general video-understanding system. It does not generate written answers, hold a conversation, write code, search the web, or interpret unrestricted video content. Instead, it performs gaze estimation and eye-gaze redirection on a defined eye-region input.
NVIDIA makes Eye Contact available as a downloadable NVIDIA NIM microservice, with documentation and a demonstration through NVIDIA Build. NIM is NVIDIA's packaging and serving approach for deploying AI inference services. In this case, the service can be hosted or self-managed on compatible NVIDIA hardware.
What the model does
At a high level, Eye Contact receives an image of the eye region and a requested redirection angle. It estimates the existing gaze, modifies the relevant internal representation, and generates a new eye patch in which the gaze has been redirected. The result is intended to be composited back into the surrounding face and video frame by the application using it.
The model uses an encoder-decoder architecture. The encoder converts the eye image into latent representations, meaning compact internal features that describe information such as the subject, head pose, environmental conditions, and gaze direction. A transformation is applied to the gaze-related information, and the decoder produces the redirected eye region.
This design makes the model useful for a focused image transformation, but it also defines its boundaries. An application must generally manage the surrounding video pipeline, including face and eye detection, frame extraction, eye-patch preparation, compositing, video decoding, and output encoding.
Inputs and outputs
The documented neural-network interface uses a normalized RGB eye patch measuring 256 × 64 pixels. The input also includes a two-element angle vector containing pitch and yaw values in radians. Pitch represents vertical gaze movement, while yaw represents horizontal movement.
The direct model outputs are:
- A redirected RGB eye image.
- An estimated gaze-angle vector.
- Eye landmarks represented as fiducial points.
At the NIM service level, NVIDIA documents MP4 video encoded with H.264 as an input format. The service can return processed video, but this should not be confused with the neural network's direct output: the underlying model produces an eye patch and related gaze and landmark data, while the service can wrap that operation in a broader video workflow.
The supplied specifications do not identify a conventional context window, maximum token count, maximum generated-text length, or other language-model output limit. Those limits are not applicable to Eye Contact's direct image-processing interface.
Where it fits in NVIDIA's catalog
Eye Contact belongs to NVIDIA Maxine, a family of technologies for audio, video, and communication-enhancement workloads. Within that family, Eye Contact occupies a specific gaze-correction role rather than serving as a general-purpose generative model.
Its current documented distribution is through NVIDIA NIM and an associated downloadable container. The support documentation lists NVIDIA GPU architectures including Volta, Turing, Ampere, Ada, and Blackwell. Supported configurations include professional and consumer GPUs with Tensor Cores, and NVIDIA recommends NVENC and NVDEC support for video workloads.
This hardware orientation is important when evaluating the model. Eye Contact is not primarily a consumer application that hides infrastructure requirements behind a simple account. A deployment normally requires compatible NVIDIA hardware, the appropriate software environment, and an application capable of preparing and processing the video stream.
Strengths and practical benefits
- Focused output: The model is purpose-built for gaze correction, so its output corresponds directly to a common video-conferencing problem.
- Useful technical outputs: In addition to the redirected eye patch, it provides estimated gaze angles and eye landmarks that can support a larger processing pipeline.
- Video-service integration: The NIM interface supports MP4/H.264 video processing, reducing the amount of service infrastructure an application must build around the model.
- Deployable on NVIDIA hardware: The documented support across several NVIDIA GPU generations provides options for workstation, professional, and data-center deployment, subject to the relevant compatibility requirements.
- Suitable for real-time-oriented workflows: NVIDIA positions the model for video conferencing and related communication applications. Actual performance will depend on the GPU, video resolution, pipeline design, and deployment configuration; the supplied research does not provide a universal latency benchmark.
Its main advantage is specialization. An organization needing a gaze-redirection component does not need to use a large vision-language model for a task that can be expressed as a constrained image transformation.
Limitations and trade-offs
Eye Contact is not a general-purpose AI assistant. It has no documented text-generation output, conversational reasoning, web search, function calling, coding capability, speech processing, image generation, or general video-understanding capability. The model's reasoning and coding scores in the supplied catalog are editorial database fields, not provider-published measures, and they should not be interpreted as evidence that the model performs those tasks.
The input format is also constrained. The documented model expects a correctly prepared eye patch and a pitch/yaw instruction rather than an arbitrary image or full video frame. A production application must locate the face and eyes, crop the appropriate region, normalize the image, apply the model output to the original frame, and handle failures such as poor lighting, occlusion, unusual poses, or tracking errors. The supplied documentation does not establish a universal quality threshold for those conditions.
The model's hardware requirements can make it less convenient than a cloud video API or a software solution designed for broader device compatibility. Conversely, a self-managed NVIDIA deployment may be preferable where an organization needs control over processing location, pipeline integration, or recurring service usage. No public token-based or per-minute price was identified for the model, so total cost depends on hardware, hosting, licensing, and implementation choices.
Pricing and availability
No public model price, token price, or recurring subscription price was identified in the supplied NVIDIA sources. Eye Contact is described as a downloadable NIM model or container rather than as a conventional consumer subscription with a simple monthly tier.
NVIDIA's model and demonstration materials apply license terms that govern evaluation and deployment. Organizations should review the applicable NVIDIA Maxine, NIM, container, and hardware requirements before using the model commercially. The absence of a listed API price should not be treated as evidence that production deployment is cost-free.
Supported modalities and capabilities
| Capability | Supported or documented behavior |
|---|---|
| Image input | Yes; normalized RGB eye patch, documented at 256 × 64 pixels |
| Video input | Yes at the NIM service level; MP4/H.264 is documented |
| Direct image output | Yes; redirected RGB eye patch |
| Video output | Available at the service level when processing video |
| Text input or output | Not documented for the model |
| Audio input or output | Not documented for the model |
| Reasoning | Not a general reasoning model |
| Coding | Not supported as a model task |
| Tools or function calling | Not documented |
| Structured text or JSON generation | Not documented as a model capability |
The model's multimodal nature should therefore be understood as visual input and visual output, not as broad multimodal conversation. It can work with video through the surrounding NIM service, but it does not analyze video in the way a general vision-language model would.
When to choose NVIDIA Eye Contact
Choose Eye Contact when the primary requirement is gaze correction and the deployment environment can support NVIDIA GPUs. It is a good fit for:
- Video-conferencing software that wants to make screen-based eye contact appear more natural.
- Recorded presentations where the speaker's gaze should be redirected toward the viewer.
- Telepresence and digital-human systems that need a dedicated eye-contact component.
- Media-processing pipelines that already use NVIDIA hardware and need an on-premises or self-managed inference service.
A specialized model is often more practical than a large generative model when the task is limited to one controlled transformation. It can reduce unnecessary model capability and keep the processing objective clear. However, the supplied research does not provide enough benchmark data to claim that it is faster or cheaper than every alternative. Those comparisons depend on hardware, resolution, frame rate, batching, and the amount of surrounding video processing.
When another option may be more appropriate
Use a general vision-language model when the application must describe video, answer questions about scenes, extract information from varied footage, or combine visual input with natural-language reasoning. Use a speech or avatar platform when the main requirement is voice interaction, lip synchronization, or full digital-human generation rather than gaze correction alone.
A cloud-based video-processing service may be easier for teams that do not want to operate NVIDIA infrastructure. A broader facial-animation or media-effects solution may also be preferable when the application needs many transformations beyond the eyes. These alternatives may trade away the focused deployment model of Eye Contact for simpler integration or wider functionality.
Technical summary
| Property | Verified information |
|---|---|
| Provider | NVIDIA |
| Model family | NVIDIA Maxine Eye Contact |
| Model ID | maxine-eye-contact |
| Primary task | Gaze estimation and eye-gaze redirection |
| Architecture | Convolutional encoder-decoder network |
| Direct input | Normalized RGB eye patch and pitch/yaw angle vector |
| Documented eye-patch size | 256 × 64 pixels |
| Direct output | Redirected RGB eye patch, estimated gaze angle, and eye landmarks |
| Service video format | MP4 encoded with H.264 |
| Inference technologies | TensorRT and Triton |
| Deployment | Downloadable NVIDIA NIM microservice/container |
| Public model pricing | Not identified |
NVIDIA Eye Contact is best understood as an infrastructure-ready component for a specific visual problem. Its value comes from producing a controlled gaze transformation and associated tracking data, not from broad generative or conversational intelligence. Teams evaluating it should start with the quality of their eye-detection and compositing pipeline, compatible GPU availability, licensing requirements, and the cost of operating the surrounding video system.

