Cosmos 3

Cosmos3-Nano

by NVIDIA AI · Current; downloadable open-weight model and available through an NVIDIA NIM endpoint

NVIDIA Cosmos3-Nano is a 16B downloadable omnimodal world foundation model that accepts text, images, video, audio, and action trajectories. It generates images, video, sound associated with video, text, and action predictions for Physical AI applications including robotics, autonomous driving, simulation, and synthetic-data generation. The model is available through NVIDIA's software stack and NIM endpoint, but its hardware requirements and public pricing limits should be evaluated before deployment.

Text Image generation Video generation Reasoning
NVIDIA Cosmos3-Nano is the 16B Nano model in the Cosmos 3 family of world foundation models. It is designed to model how physical environments change over time rather than serve as a general-purpose chatbot. The model supports multimodal generation and action reasoning for robotics, autonomous driving, simulation, and other Physical AI applications.
Outputs

What Cosmos3-Nano can produce

Text Image generation Video generation Audio Actions
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
2/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Cosmos 3
Model type Multimodal
Release date 2026-05-31
Status Current; downloadable open-weight model and available through an NVIDIA NIM endpoint
Knowledge cutoff notes

NVIDIA's available model documentation does not specify a knowledge cutoff for Cosmos3-Nano. The model is primarily a world-generation and physical-simulation model rather than a conventional knowledge-grounded language model.

Model notes

Cosmos3-Nano is NVIDIA's 16B model in the Cosmos 3 family. It uses a Mixture-of-Transformers architecture with autoregressive and diffusion components. Documented modes include text-to-image, text-to-video, image-to-video, video-to-video, forward dynamics, policy, and inverse dynamics. Outputs can include images, video, sound associated with video generation, text, and action predictions. NVIDIA recommends data-center or workstation hardware such as RTX PRO 6000, H100, or B200. The downloadable repository is approximately 35 GB and is released under OpenMDW1.1. The hosted NIM endpoint is presented as a free or trial endpoint, but no standard public model-specific pricing schedule was verified. Editorial scores reflect comparison with current AI models and are not NVIDIA benchmark ratings.

Model guide

NVIDIA Cosmos3-Nano: A 16B World Model for Physical AI

NVIDIA Cosmos3-Nano is a downloadable 16-billion-parameter omnimodal world foundation model for Physical AI. It accepts text, images, video, audio, and action trajectories, then generates outputs such as images, video, sound associated with video, text, and action predictions for robotics, autonomous vehicles, simulation, and synthetic-data workflows.

What is NVIDIA Cosmos3-Nano?

NVIDIA Cosmos3-Nano is a 16-billion-parameter omnimodal world foundation model from NVIDIA. Its purpose is to represent and generate possible states of the physical world. In practical terms, it can use combinations of text, images, video, audio, and action trajectories to create visual or other predictions about what may happen next.

This makes Cosmos3-Nano different from a conventional language model intended mainly for conversation, summarization, or code generation. Its documented applications include robotics research, autonomous-vehicle development, industrial simulation, smart-space scenarios, synthetic training data, and other workflows in which an AI system needs to reason about motion, environments, and actions.

NVIDIA positions the model as the Nano tier in its Cosmos 3 family. It sits between the smaller 4B Cosmos3-Edge and the larger 64B Cosmos3-Super in the supplied model positioning. Compared with the larger sibling, Nano is intended to offer a more practical balance between generation quality and deployment cost, while offering broader multimodal generation than the smaller model.

Inputs and outputs supported by the model

Cosmos3-Nano accepts several kinds of input. These include text prompts, still images, video, audio, and action trajectories. An action trajectory is a sequence describing movement or control actions, such as the path a robot or vehicle takes. Combining this information allows the model to connect observations with possible future actions and states.

The documented generation modes include:

  • Text-to-image: generating an image from a written description.
  • Text-to-video: generating a video sequence from text.
  • Image-to-video: animating or extending an image into a video sequence.
  • Video-to-video: transforming or continuing an existing video.
  • Forward dynamics: predicting future visual states from an environment and an action trajectory.
  • Policy: producing action-related predictions for a physical-world scenario.
  • Inverse dynamics: inferring actions associated with observed motion or state changes.

Its outputs can include text, images, video, sound associated with video generation, and action predictions. The supplied specifications classify its input as multimodal and its output as multimodal. Audio is supported as an input and as an output associated with video generation, but Cosmos3-Nano is not described as a standalone speech-recognition or text-to-speech model.

How Cosmos3-Nano models physical situations

The model is particularly useful when the desired result is a plausible future state rather than a factual written answer. For example, an image of a scene combined with an action trajectory can be used to predict subsequent frames. Video and action information can also support inverse-dynamics tasks, where the system estimates which actions correspond to observed movement.

This capability is relevant to synthetic-data pipelines. A robotics team could use generated scenes and future states to expand training material for perception or control systems. An autonomous-driving developer could explore simulated situations that are difficult, expensive, or unsafe to capture repeatedly in the real world. Researchers can also use generated sequences to test how downstream systems respond to changing environments.

These outputs should be treated as model-generated predictions, not guaranteed physical simulations. The supplied research does not provide benchmark results or a guarantee of physical accuracy. Any generated data used for robotics or autonomous systems should be checked for realism and suitability before it influences safety-critical decisions.

Architecture, size, and deployment requirements

Cosmos3-Nano uses NVIDIA's Cosmos 3 Mixture-of-Transformers architecture. In broad terms, the architecture combines different transformer components for different types of generation. An autoregressive transformer handles discrete token generation, while a diffusion transformer synthesizes continuous modalities such as images, video, audio, and actions through iterative denoising.

The model is available as downloadable NVIDIA weights under the OpenMDW1.1 license. The repository contains approximately 35 GB of model artifacts, so local use requires substantial storage and suitable GPU infrastructure. NVIDIA documentation lists the RTX PRO 6000, H100, and B200 among recommended hardware for the Nano model. These requirements make it a workstation, data-center, or specialized research deployment rather than a model intended for an ordinary laptop.

Cosmos3-Nano can be deployed with NVIDIA's Cosmos software stack and is also documented for vLLM-Omni serving. The downloadable model gives teams more control over deployment and data handling, but local operation also creates responsibility for hardware provisioning, software setup, performance tuning, and operational maintenance.

Hosted endpoint and pricing

NVIDIA provides a hosted Cosmos3-Nano NIM endpoint through its model catalog. The documented endpoint supports modes including text-to-video, image-to-video, video-to-video, forward dynamics, policy, and inverse dynamics. Requests are mode-specific and can reference image or video inputs through encoded media or accessible HTTP(S) URLs.

The hosted service applies input-media screening and visual SynthID watermarking. NVIDIA presents the endpoint as a downloadable free endpoint or trial service in the supplied documentation. However, no standard public per-token, per-image, or per-video price was verified for Cosmos3-Nano. The research also does not specify a universal quota, latency commitment, or maximum output duration.

Consequently, the hosted endpoint should not be evaluated as though it had a simple, confirmed consumer subscription price. Production teams should verify current access conditions, quotas, infrastructure charges, and service terms directly with NVIDIA. The downloadable weights and the hosted endpoint also involve different trade-offs: local deployment may offer greater control but requires expensive hardware, while hosted access reduces infrastructure work but depends on endpoint availability and usage conditions.

Important limits and unsupported roles

NVIDIA's available documentation does not specify a conventional context window, maximum token output, or knowledge cutoff for Cosmos3-Nano. Those omissions are meaningful because this is primarily a world-generation and physical-simulation model, not a standard knowledge-grounded language model. Users should not assume that it has the same long-context or text-generation limits as a general-purpose large language model.

The model is also not presented as a general-purpose conversational assistant. It is a poor fit for ordinary chat, broad factual question answering, software coding, embeddings, transcription, or standalone speech synthesis. Its documented tool-use setting is 0, and no general function-calling capability is specified. Structured JSON output is not verified either.

Its multimodal abilities should therefore be understood in the context of physical-world generation. Supporting video or audio does not mean that the model replaces a dedicated video editor, speech model, perception system, or robotics controller. Action predictions may be useful as part of a larger system, but they should not be treated as automatically safe control commands.

Reasoning, coding, speed, and cost trade-offs

Cosmos3-Nano's reasoning is specialized rather than general. Its useful form of reasoning concerns relationships among observations, motion, actions, and future physical states. The editorial evaluation supplied for this model assigns a reasoning score of 7 out of 10 and a coding score of 2 out of 10. These are editorial comparison scores, not NVIDIA-published benchmark results.

The low coding assessment reflects the model's intended role: it is designed for Physical AI and world generation, not software development. A separate language or code model is more appropriate for writing programs, debugging, generating documentation, or operating as a conversational agent.

The supplied editorial scores assign Cosmos3-Nano a speed score of 6 and a cost score of 8. These scores are also subjective evaluations rather than provider specifications. The cost advantage is relative to larger world models such as Cosmos3-Super, not evidence that local deployment is inexpensive. A 16B multimodal model with approximately 35 GB of artifacts still requires serious storage and GPU resources. Nano's practical advantage is that it should be less demanding than the 64B sibling while retaining a broad set of generation modes.

When to choose Cosmos3-Nano

Cosmos3-Nano is a strong candidate when the central problem involves physical environments, multimodal world generation, or action-conditioned prediction. It is particularly suitable for:

  • Generating synthetic video and visual scenarios for Physical AI research.
  • Predicting future frames from an observed scene and planned actions.
  • Exploring robotics behaviors and action-conditioned environments.
  • Creating autonomous-vehicle simulation scenarios and training data.
  • Studying inverse dynamics from video and other observations.
  • Building research systems that need downloadable weights rather than only a hosted interface.

Choose it over a general-purpose language model when physical-world state, video generation, or action reasoning is more important than conversation and coding. Choose it over the larger Cosmos3-Super when deployment resources and operating cost are more constrained and the additional scale of the 64B model is not necessary. The smaller Cosmos3-Edge may be more appropriate when minimizing hardware demands is the overriding priority, although the supplied research indicates that Nano offers broader multimodal generation than that smaller tier.

Another option may be better when the task is ordinary chat, code generation, document analysis, web research, speech transcription, standalone speech synthesis, or embeddings. Cosmos3-Nano should also be combined with conventional perception, planning, validation, and safety systems rather than used alone to control a real robot or vehicle.

Overall assessment

NVIDIA Cosmos3-Nano is best understood as a specialized 16B world model for Physical AI. Its distinguishing feature is not general language intelligence but the ability to condition generation on multiple forms of physical-world information and produce future visual, audio-associated, textual, or action-related outputs.

Its downloadable weights, NIM endpoint, and support for video, images, audio, and action trajectories give it a broad deployment profile. At the same time, its hardware requirements, unspecified context and output limits, lack of verified public model-specific pricing, and limited suitability for chat or coding make careful project selection essential. For robotics, simulation, autonomous-vehicle research, and synthetic-data workflows, those trade-offs may be worthwhile; for general-purpose AI assistance, a different model type is more appropriate.


Answers to Frequently Asked Questions

How does Cosmos3-Nano compare with Cosmos3-Edge and Cosmos3-Super?
Cosmos3-Nano is positioned between the smaller 4B Cosmos3-Edge and the larger 64B Cosmos3-Super. It aims to provide broader multimodal generation than Cosmos3-Edge while requiring fewer resources than Cosmos3-Super. It may be a practical choice when teams need extensive world-generation capabilities but have more limited deployment resources.
Is Cosmos3-Nano suitable for general chat, coding, or speech recognition?
No. Cosmos3-Nano is specialized for physical-world generation, multimodal prediction, and action-related reasoning rather than general conversation, software coding, embeddings, transcription, or standalone speech synthesis. A general-purpose language, code, perception, or speech model is better suited to those tasks.
What hardware and deployment options are required for Cosmos3-Nano?
The downloadable model contains approximately 35 GB of artifacts and requires substantial storage and GPU capacity. NVIDIA lists the RTX PRO 6000, H100, and B200 among recommended hardware. Cosmos3-Nano can be deployed locally through NVIDIA's Cosmos software stack or served with vLLM-Omni, and NVIDIA also provides a hosted Cosmos3-Nano NIM endpoint.
What inputs and outputs does Cosmos3-Nano support?
Cosmos3-Nano can accept text, images, video, audio, and action trajectories. Its documented capabilities include text-to-image, text-to-video, image-to-video, video-to-video, forward dynamics, policy prediction, and inverse dynamics. Outputs can include text, images, video, audio associated with video generation, and action-related predictions.
What is NVIDIA Cosmos3-Nano used for?
NVIDIA Cosmos3-Nano is a 16-billion-parameter omnimodal world model designed for Physical AI applications. It can support robotics research, autonomous-vehicle development, industrial simulation, smart-space scenarios, synthetic training data, and workflows that require reasoning about environments, motion, actions, and possible future states.


Sources 6
Provider

About NVIDIA AI