Cosmos 3

Cosmos3-Edge

by NVIDIA AI · Current; open model; gated Hugging Face access

NVIDIA Cosmos3-Edge is an approximately 4-billion-parameter multimodal world model for physical AI. It supports text, image, video, and action inputs, produces text, image, video, and action-oriented outputs, and targets low-latency deployment on NVIDIA edge and GPU platforms. Its main trade-offs are specialized hardware requirements, no current audio-generation path, and no documented hosted token pricing.

Text Image generation Video generation Reasoning
NVIDIA Cosmos3-Edge is the compact Edge model in NVIDIA's Cosmos 3 world-model family. It is designed for robots, autonomous systems, and visual agents that need to interpret physical environments, reason about possible changes, and generate visual or action-oriented outputs without relying exclusively on large data-center infrastructure. The open checkpoint is approximately 4 billion parameters and is available through a gated NVIDIA repository on Hugging Face.
Outputs

What Cosmos3-Edge can produce

Text Image generation Video generation Actions
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
3/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Cosmos 3
Model type Multimodal
Context window 131K tokens
Release date 2026-07-20
Status Current; open model; gated Hugging Face access
Knowledge cutoff notes

NVIDIA does not publish a direct knowledge-cutoff date for the exact Cosmos3-Edge checkpoint in the authoritative model documentation reviewed. The model is primarily a physical-world vision, generation, and action model rather than a conventional continuously updated knowledge model.

Model notes

Cosmos3-Edge is NVIDIA's approximately 4-billion-parameter Edge member of the Cosmos 3 world-model family. NVIDIA documentation identifies it for edge or on-device deployment on platforms including Jetson AGX Orin, Jetson Thor, and RTX PRO GPUs. The model accepts text, image, video, and action inputs and supports text, image, video, and action-oriented outputs. The maintained Cosmos Framework documentation states that the current Edge checkpoint supports all listed modes except audio because it does not include a sound tokenizer. The published configuration specifies a 131,072-token maximum positional length for the text component. NVIDIA provides fine-tuning recipes for vision generation and physical-plausibility alignment. No official hosted API token pricing was found; the model is distributed as an open checkpoint through NVIDIA's gated Hugging Face repository.

Model guide

NVIDIA Cosmos3-Edge: An Open 4B World Model for Real-Time Physical AI

NVIDIA Cosmos3-Edge is a compact, approximately 4-billion-parameter multimodal world model for physical AI. It accepts text, images, video, and action data, and can produce text, images, video, and action sequences. Its main distinction is an emphasis on memory-efficient, low-latency inference for robotics, autonomous systems, visual reasoning, and world simulation on edge-oriented NVIDIA hardware.

What is NVIDIA Cosmos3-Edge?

NVIDIA Cosmos3-Edge is a multimodal world model for physical AI. In practical terms, it is designed to help an AI system understand changing physical environments rather than function primarily as a general-purpose chat or coding assistant. It can combine language with visual information and action data to support tasks such as robotic perception, visual reasoning, policy development, and synthetic world generation.

The model is the compact Edge member of NVIDIA's Cosmos 3 family. Its approximately 4-billion-parameter size is intended to make deployment more feasible on edge and workstation hardware than deployment of a much larger world model. NVIDIA identifies platforms including Jetson AGX Orin, Jetson Thor, RTX PRO systems, and other compatible NVIDIA GPU environments as target deployment options.

Cosmos3-Edge is distributed as an open checkpoint under NVIDIA's model licensing terms, but access to the Hugging Face repository is gated. Users must request access and authenticate before downloading it. This is different from a conventional hosted chatbot: the supplied research does not identify an official hosted per-token API or a standard consumer subscription for this model.

Purpose and position in the Cosmos 3 family

Cosmos3-Edge is positioned for situations where inference speed, memory efficiency, and local deployment matter. The model targets physical-AI applications that may need to respond close to the source of sensor data, including robots, factory systems, autonomous machines, and edge video-analysis systems.

Its smaller scale comes with a trade-off. NVIDIA's larger Cosmos3-Nano and Cosmos3-Super variants are more appropriate when greater model capacity, higher visual quality, broader generation support, or large-scale data-center execution is the priority. Cosmos3-Edge instead emphasizes practical deployment on comparatively constrained NVIDIA hardware.

This positioning makes it different from a general-purpose language model. It is not presented as an enterprise chat model, coding model, or broad knowledge assistant. Its value comes from connecting visual understanding, physical reasoning, world generation, and action prediction in one physical-AI-oriented checkpoint.

Inputs and outputs

Cosmos3-Edge supports several input types relevant to physical environments:

  • Text: prompts and textual reasoning tasks.
  • Images: visual observations and image-conditioned workflows.
  • Video: temporal visual information for understanding and generation tasks.
  • Action data: action conditioning and robot-policy-oriented workflows.

Depending on the integration and selected mode, it can produce text, images, videos, and action sequences. Action output is particularly relevant to robotics and embodied agents, where a model may need to predict or support a sequence of operations rather than return only a textual answer.

The current Edge checkpoint does not support the audio-enabled generation path described for some larger Cosmos 3 variants. NVIDIA's maintained Cosmos Framework documentation attributes this limitation to the absence of the required sound tokenizer. Audio input and audio output should therefore not be treated as supported features for this checkpoint.

Video-to-video transfer is also not currently supported in the maintained Cosmos 3 model reference. Published Edge-specific generation constraints include 256p and 480p resolutions, frame rates from 12 to 30 frames per second, and generation lengths of approximately 50 to 150 frames. These are documented workflow constraints, not a guarantee that every integration exposes exactly the same settings.

Architecture and context length

The published configuration describes a multimodal architecture with a vision encoder, a projector, and a dense language backbone. The text component uses a 28-layer decoder with a 2,048-dimensional hidden size. Its vision component is based on SigLIP2, a vision encoder architecture used to convert image information into representations that the language and generation components can process.

The published text configuration specifies a maximum positional length of 131,072 tokens. This is the model's documented context-related positional limit for the text component. It should not be interpreted as a promise that every image or video workflow can freely consume an equivalent amount of practical input data: visual encodings, preprocessing, memory limits, and the selected framework can affect usable capacity.

Transformers integration exposes the reasoner portion for text responses from text, image, or video inputs. Diffusers and related Cosmos integrations provide generator-oriented workflows for visual generation and action-related modes. The exact software path therefore depends on whether the task is reasoning over observations, generating visual content, or supporting a physical-AI policy workflow.

Deployment, speed, and cost trade-offs

Cosmos3-Edge is intended for low-latency inference on NVIDIA edge, workstation, and GPU platforms. Running a model locally can reduce dependence on a remote inference service and may be useful where sensor data must remain near the robot or machine. It also gives developers more control over runtime configuration and integration with a physical system.

The trade-off is operational complexity. Users need compatible NVIDIA hardware, a suitable Linux-oriented software stack, drivers, model files, and enough memory for the chosen workflow. The model is not a turnkey web application. Hardware requirements and real-time performance will vary with resolution, frame rate, sequence length, quantization or optimization choices, and the complexity of the generation pipeline.

There is no official hosted price per input or output token identified in the supplied research. The checkpoint itself is available through a gated model repository, but local deployment still has infrastructure costs, including GPU hardware, power, storage, engineering time, and maintenance. For teams without suitable NVIDIA hardware, a hosted model or a larger cloud-oriented service may be simpler even if its recurring usage cost is more visible.

Reasoning, coding, and tool support

Cosmos3-Edge is built for physical reasoning: interpreting visual scenes, modeling possible changes in an environment, and supporting action or policy-oriented inference. This makes its reasoning useful for embodied applications, but it should not be confused with a general-purpose reasoning model optimized for mathematics, research, or long-form business analysis.

It is not intended as a coding model. The supplied research does not document specialized code-generation capabilities, software-agent behavior, function calling, or external tool execution. Developers can integrate the checkpoint into a larger robotics or computer-vision system, but that surrounding orchestration should not be attributed to Cosmos3-Edge itself.

Similarly, no verified streaming, batch API, JSON mode, caching, or hosted function-tool interface is specified for the open checkpoint. Structured action sequences may be part of a model workflow, but that is not the same as a documented general-purpose JSON or tool-calling API.

Fine-tuning and customization

NVIDIA provides supervised fine-tuning recipes for Cosmos3-Edge. The documented workflows include vision-generator fine-tuning and physical-plausibility alignment using VideoPhy-2. These options are relevant when a developer needs the model to reflect a particular environment, video domain, robotics dataset, or physical behavior profile.

Fine-tuning can make the checkpoint more useful for a specialized deployment, but it also increases the engineering burden. Teams need suitable training data, compatible GPU capacity, evaluation procedures, and a way to check that generated actions or scenes remain physically plausible. Fine-tuning should not be assumed to add unsupported modalities such as audio to the current Edge checkpoint.

Main strengths and limitations

Strengths

  • Compact physical-AI focus: its approximately 4-billion-parameter scale is aimed at more practical edge and workstation deployment than larger world-model variants.
  • Broad multimodal coverage: it combines text, image, video, and action inputs with text, image, video, and action-oriented outputs.
  • Edge-oriented design: it is intended for low-latency inference on platforms such as Jetson systems and NVIDIA GPUs.
  • Useful robotics integration: action conditioning and action-sequence prediction support policy-oriented experimentation.
  • Customization options: NVIDIA provides fine-tuning recipes for visual generation and physical-plausibility alignment.

Limitations

  • No audio generation path: the current Edge checkpoint does not include the sound tokenizer required for the documented audio-enabled workflow.
  • Hardware dependence: practical use requires compatible NVIDIA GPU hardware and a suitable software environment.
  • Lower capacity than larger siblings: the compact design may be less suitable when maximum visual quality or large-scale generation is more important than latency and deployment size.
  • No documented hosted pricing: users should plan for local infrastructure rather than assuming a ready-made pay-per-token endpoint.
  • Not a general chat or coding model: its physical-world specialization is an advantage for robotics but a limitation for ordinary conversational, coding, or enterprise knowledge tasks.
  • Workflow-specific constraints: supported resolutions, frame rates, frame counts, and generation modes depend on the maintained Cosmos integrations.

When to choose Cosmos3-Edge

Choose Cosmos3-Edge when the primary problem involves physical environments and the deployment needs to stay relatively close to the edge. It is a strong candidate for robotic perception and policy prototyping, visual reasoning over camera data, autonomous-system experiments, synthetic visual-data generation, and world-model research on Jetson or NVIDIA GPU hardware.

Its compact scale is especially relevant when the system needs lower latency or a smaller deployment footprint than a larger world model can provide. Local execution may also be preferable when sending camera or sensor data to a remote service is undesirable or impractical.

Another option may be more appropriate when the task is general-purpose chat, code generation, web research, audio generation, or large-scale high-quality video production. A larger Cosmos variant is a better fit when the priority is model capacity, visual quality, or broader generation support. A hosted general-purpose model may be easier for teams that do not have compatible NVIDIA hardware or do not want to operate an on-device inference stack.

Overall, Cosmos3-Edge is best understood as a compact physical-AI foundation model rather than a universal assistant. Its central trade-off is clear: it gives developers multimodal world understanding and action-oriented workflows in a smaller edge-focused package, while requiring specialized hardware and accepting lower capacity and narrower general-purpose coverage than larger or cloud-oriented alternatives.


Answers to Frequently Asked Questions

When should developers choose NVIDIA Cosmos3-Edge?
Developers should consider Cosmos3-Edge when they need a relatively compact world model for robotics, autonomous systems, camera-based visual reasoning, synthetic data generation, or physical-AI research on NVIDIA hardware. A larger Cosmos variant or a hosted general-purpose model may be more suitable for maximum visual quality, audio generation, large-scale video production, or teams without compatible NVIDIA infrastructure.
Is NVIDIA Cosmos3-Edge a general-purpose chat or coding model?
No. Cosmos3-Edge is specialized for physical AI, visual understanding, world modeling, and action-oriented workflows. It is not presented as a general-purpose chat, coding, research, or enterprise knowledge model, and the supplied research does not document specialized code generation, function calling, or external tool execution.
What hardware is needed to run NVIDIA Cosmos3-Edge?
NVIDIA Cosmos3-Edge is intended for compatible NVIDIA edge, workstation, and GPU platforms, including Jetson AGX Orin, Jetson Thor, and RTX PRO systems. Actual performance and hardware requirements depend on factors such as resolution, frame rate, sequence length, optimization, and the selected generation workflow.
What is NVIDIA Cosmos3-Edge designed for?
NVIDIA Cosmos3-Edge is a multimodal world model for physical AI. It is designed for robotic perception, visual reasoning, policy development, synthetic world generation, and other applications that require AI to understand changing physical environments.
What inputs and outputs does NVIDIA Cosmos3-Edge support?
Cosmos3-Edge supports text, images, video, and action data as inputs. Depending on the workflow, it can produce text, images, videos, and action sequences. The current Edge checkpoint does not support the documented audio-enabled generation path or video-to-video transfer.


Sources 9
Provider

About NVIDIA AI