Cosmos 3

Cosmos3-Super

by NVIDIA AI · Current; open-weight model; commercially and non-commercially usable

NVIDIA Cosmos3-Super is a 64-billion-parameter open omnimodal world model for Physical AI. It combines visual and temporal reasoning with generation of images, video with synchronized audio, text, and embodiment-specific action trajectories for robotics, autonomous vehicles, simulation, and synthetic-data workflows. Its main trade-offs are substantial multi-GPU requirements, no identified official per-token price, and limited suitability for ordinary chat, coding, or edge deployment.

Text Image generation Video generation Reasoning
NVIDIA Cosmos3-Super is designed to help machines understand and predict physical environments rather than simply generate text. Its Mixture-of-Transformers architecture combines an autoregressive reasoner with a diffusion-based generator, allowing the model to interpret text, images, video, and action information and then produce simulated visual futures, audio-containing video, and compatible action trajectories. The model is intended for research and production teams with access to substantial data-center GPU resources, not for lightweight consumer chat or ordinary coding assistance.
Outputs

What Cosmos3-Super can produce

Text Image generation Video generation Audio Actions
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

8/10 Reasoning
3/10 Coding
2/10 Speed
3/10 Cost efficiency
Specifications

Technical details

Model family Cosmos 3
Model type Multimodal
Context window 262K tokens
Release date 2026-05-31
Status Current; open-weight model; commercially and non-commercially usable
Knowledge cutoff notes

NVIDIA does not publish a conventional textual knowledge-cutoff date for Cosmos3-Super. The model is primarily a Physical AI world model trained for multimodal reasoning and generation rather than a standard factual chat model.

Model notes

Cosmos3-Super is a 64B Mixture-of-Transformers model with separate reasoning and generation workflows. The reasoner supports text, image, and video inputs with up to 256K tokens of documented long-context input. The generator supports text, images, video, audio-containing video, and action trajectories. Generated audio is stereo AAC at 48 kHz muxed into MP4 video rather than exposed as standalone audio. Action generation is limited to documented compatible embodiments and dimensionalities. NVIDIA recommends H200, B200, or GB200 data-center GPUs. The model is distributed as open weights under OpenMDW 1.1; no official NVIDIA per-token API price was identified for this exact checkpoint.

Model guide

NVIDIA Cosmos3-Super: Open 64B World Model for Physical AI

NVIDIA Cosmos3-Super is a 64-billion-parameter open omnimodal world model for Physical AI. It combines multimodal reasoning with generation of images, video with synchronized audio, text, and embodiment-specific action trajectories for robotics, autonomous vehicles, simulation, and synthetic-data creation.

What is NVIDIA Cosmos3-Super?

NVIDIA Cosmos3-Super is a 64-billion-parameter open omnimodal world model for Physical AI. In practical terms, it is built to represent and generate situations in the physical world: a robot moving through a workspace, a vehicle traveling through a scene, or an environment changing over time.

The model is aimed at robotics, autonomous vehicles, industrial systems, smart spaces, simulation, and synthetic-data generation. Rather than functioning primarily as a general-purpose conversational model, Cosmos3-Super can use visual and action information to model possible physical-world outcomes. It is also positioned as a teacher model for distillation and post-training, where its outputs can help train smaller or more specialized systems.

NVIDIA distributes the model as open weights under the OpenMDW 1.1 license for commercial and non-commercial use, according to the supplied model information. Open weights can give research and engineering teams more control than a closed hosted service, but running a model of this size still requires considerable infrastructure.

How the model is built

Cosmos3-Super uses a Mixture-of-Transformers architecture with two complementary processing paths. An autoregressive transformer handles discrete token generation and reasoning. A diffusion transformer generates continuous modalities such as images, video, sound, and actions.

This separation supports two related workflows. The reasoner interprets multimodal inputs and can return textual reasoning, visual grounding information, temporal understanding, point localization, and bounding-box coordinates. The generator then produces simulated or predicted content, including visual scenes, audio-containing video, and action trajectories.

For a beginner, the distinction is useful: the reasoner is closer to an analysis and understanding component, while the generator is intended to create possible future states or training examples. The model is therefore better understood as a world-modeling system than as a conventional text-only large language model.

Supported inputs and outputs

The documented reasoner accepts text, text with images, and text with video. Its documented context capacity is up to 256K tokens. This long context can be useful when a task requires extended multimodal information, although the supplied research does not specify a universal maximum output size.

The generator accepts several input types:

  • Text prompts
  • JPG, PNG, JPEG, or WebP images
  • MP4 video
  • JSON action trajectories

Images and videos can use documented resolutions of 256p, 480p, and 720p. Common supported aspect ratios include 16:9, 4:3, 1:1, 3:4, and 9:16. Video inputs may include audio; NVIDIA documents stereo audio at 48 kHz for muxed MP4 inputs.

Outputs can include text, JPEG images, MP4 video, and JSON-compatible one-dimensional action lists. Generated video can contain synchronized audio encoded as stereo AAC at 48 kHz and muxed into the MP4 file. This is different from offering a separate audio-generation endpoint: the supplied documentation describes audio as part of generated video rather than as a standalone audio file.

Documented video generation supports durations from five to 400 frames, with 189 frames listed as a default generation duration. Supported frame rates include 10, 16, 24, and 30 frames per second, subject to the selected runtime and checkpoint configuration.

What Cosmos3-Super is used for

The model's main value is its ability to generate plausible physical-world futures and multimodal training data. A robotics team could condition generation on an image, video, text description, or action sequence and use the resulting scenes for simulation or post-training. An autonomous-driving team could use it to explore future visual states under documented vehicle-motion conditions.

Potentially suitable workloads include:

  • Generating synthetic image and video data for perception or embodied-AI training
  • Predicting or simulating future visual scenes from current observations
  • Creating teacher-model outputs for distillation into smaller models
  • Researching robot behavior and action-conditioned world models
  • Supporting autonomous-vehicle and industrial simulation workflows
  • Grounding objects and locations in images or video through points and bounding boxes

Action generation is embodiment-specific rather than universally compatible with any robot. Documented examples include camera motion, autonomous-vehicle motion, egocentric motion, single- and dual-arm robots, AgiBot platforms, UR robots, Google robots, WidowX 250, and UMI configurations. Generated actions therefore need to be mapped to a supported embodiment and validated before they are used in simulation or hardware.

Reasoning, coding, and tool support

Cosmos3-Super has meaningful multimodal reasoning capabilities, especially for visual grounding, temporal understanding, point localization, and bounding-box coordinates. Its reasoner can work with text, images, and video, making it appropriate for interpreting physical scenes and events over time.

It is not primarily a coding model. The supplied evaluation records a low coding suitability score, and the model's purpose is world modeling rather than software development. It may be useful for producing structured action information or supporting a larger Physical AI pipeline, but teams seeking code generation, debugging, or conventional programming assistance should select a model designed for those tasks.

No tool-use or function-calling capability is documented for this exact checkpoint. Action trajectories are model outputs, not evidence of general-purpose API tool use. Similarly, no verified standalone JSON mode, streaming feature, batch API, caching system, or maximum output-token limit is identified in the supplied research.

Hardware, deployment, and speed trade-offs

Cosmos3-Super is a large model intended for data-center deployment. NVIDIA lists H200, B200, and GB200 systems as recommended hardware in its model matrix. The model repository also provides multi-GPU configurations, and documented serving examples distribute the model across multiple high-memory GPUs.

It can be downloaded from NVIDIA's Hugging Face repository and deployed through NVIDIA's Cosmos software, compatible diffusion tooling, vLLM-Omni, or NVIDIA NIM-based infrastructure. These are deployment routes rather than a single guaranteed hosted API experience, so the exact setup, latency, and operating cost depend on the chosen runtime and hardware.

The principal trade-off is capability versus speed and cost. A 64-billion-parameter model that generates video, synchronized audio, and action trajectories requires substantially more compute than a small text model or an edge-focused multimodal model. The supplied evaluation rates its speed and cost suitability below its reasoning capability, reflecting this infrastructure burden. Those scores are editorial evaluations, not NVIDIA-published benchmarks.

Cosmos3-Super is consequently a better fit for organizations operating multi-GPU servers or cloud infrastructure than for users with ordinary workstations. Teams that need low-latency deployment on constrained devices should investigate smaller Cosmos models, including Cosmos3-Nano or Cosmos3-Edge, when their workload can accept reduced scale or capability.

Context limits and practical constraints

The reasoner's documented context capacity is up to 256K tokens. This is a substantial input window for multimodal reasoning, but it should not be interpreted as a promise that every workflow can process arbitrarily long or high-resolution media. Video duration, resolution, frame rate, runtime, and checkpoint configuration all affect the practical workload.

The generator supports video from five to 400 frames in the supplied documentation, but the actual computational cost rises with resolution, duration, and generation settings. The research does not identify a separate maximum output-token limit for text, so no such limit should be assumed.

Generated actions also require careful validation. A plausible trajectory is not automatically safe, physically feasible, or suitable for a particular robot. Applications should test outputs in simulation, apply task-specific constraints, and retain hardware-level safety controls. NVIDIA's default inference stack includes Cosmos Guardrail components, but guardrails do not replace application-level validation.

Pricing and availability

No official NVIDIA per-token price was identified for the Cosmos3-Super checkpoint. The model is distributed as open weights rather than presented as a conventional hosted chat model with a standard monthly plan or predictable input and output token rates.

That does not mean deployment is free. Users may incur costs for data-center GPUs, cloud instances, storage, networking, engineering, and operational support. The total cost depends on whether the model is self-hosted, deployed through an infrastructure provider, or accessed through a separate service built around NVIDIA technology. Any third-party hosting price should therefore be treated as a provider-specific service price, not as the official price of the model itself.

Main strengths and limitations

AreaAssessment
Primary strengthCombines multimodal reasoning with world simulation, video generation, synchronized audio, and action generation.
Input coverageText, images, video, audio-containing video, and JSON action trajectories are documented.
Output coverageText, images, MP4 video with muxed audio, and embodiment-specific action lists.
CustomizationOpen-weight distribution supports self-hosting and research customization.
Main limitationRequires substantial multi-GPU infrastructure and is not designed for lightweight edge use.
Action limitationAction outputs are tied to documented embodiments and must be validated before deployment.
Pricing limitationNo official per-token price is identified for this exact checkpoint.

When to choose Cosmos3-Super

Choose Cosmos3-Super when the project needs a large, open multimodal world model and can support the associated hardware and engineering requirements. It is particularly appropriate for Physical AI teams working on synthetic data, robotics research, autonomous-vehicle simulation, multimodal scene understanding, or teacher-model distillation.

Its open-weight format may also appeal to organizations that need to inspect, customize, or self-host a model rather than rely on a closed general-purpose API. The combination of a long-context reasoner and a multimodal generator is useful when the same research workflow needs both scene interpretation and future-state generation.

Another option is more appropriate when the priority is fast inference, low operating cost, consumer chat, general coding, edge deployment, or a simple managed API with published token pricing. NVIDIA's smaller Cosmos3-Nano and Cosmos3-Edge models may be better suited to constrained deployments, while a dedicated language or coding model is a better choice for ordinary text and software tasks.

Bottom line

Cosmos3-Super is a specialized open model for building and studying Physical AI systems. Its distinguishing capability is not ordinary conversation but the ability to connect multimodal understanding with generated visual futures, synchronized audio, and action trajectories. That makes it relevant to robotics, autonomous vehicles, simulation, and synthetic-data pipelines.

The same specialization creates its main drawbacks: high infrastructure requirements, embodiment-specific action support, no identified official per-token price, and limited relevance to conventional chat or coding. For teams with the necessary GPU resources and a physical-world modeling problem, Cosmos3-Super offers a broad open-weight foundation. For smaller, faster, or more predictable deployments, a lighter model or a purpose-built hosted service is likely to be more practical.


Answers to Frequently Asked Questions

Can Cosmos3-Super generate actions for any robot?
No. Its action generation is embodiment-specific and documented for selected platforms and configurations, including autonomous vehicles, single- and dual-arm robots, AgiBot platforms, UR robots, Google robots, WidowX 250, and UMI configurations. Action trajectories must be mapped to a supported embodiment, tested in simulation, and validated with application-level and hardware safety controls.
Does NVIDIA Cosmos3-Super have an official per-token price?
No official NVIDIA per-token price has been identified for the Cosmos3-Super checkpoint. It is distributed as open weights rather than as a conventional hosted chat model. Deployment can still incur costs for GPUs, cloud infrastructure, storage, networking, engineering, and operations.
What is NVIDIA Cosmos3-Super used for?
NVIDIA Cosmos3-Super is an open 64-billion-parameter world model for Physical AI. It is designed for robotics, autonomous vehicles, industrial simulation, smart spaces, synthetic-data generation, multimodal scene understanding, and teacher-model distillation.
What inputs and outputs does NVIDIA Cosmos3-Super support?
The reasoner supports text, images, and video, with a documented context capacity of up to 256K tokens. The generator accepts text, JPG, PNG, JPEG, or WebP images, MP4 video, and JSON action trajectories. Outputs can include text, JPEG images, MP4 video with synchronized audio, and JSON-compatible action lists.
What hardware is required to run NVIDIA Cosmos3-Super?
Cosmos3-Super is intended for data-center deployment and generally requires multiple high-memory GPUs. NVIDIA lists H200, B200, and GB200 systems as recommended hardware. It can be deployed through NVIDIA Cosmos software, compatible diffusion tooling, vLLM-Omni, or NVIDIA NIM-based infrastructure.


Sources 5
Provider

About NVIDIA AI