What is NVIDIA Cosmos3-Super?
NVIDIA Cosmos3-Super is a 64-billion-parameter open omnimodal world model for Physical AI. In practical terms, it is built to represent and generate situations in the physical world: a robot moving through a workspace, a vehicle traveling through a scene, or an environment changing over time.
The model is aimed at robotics, autonomous vehicles, industrial systems, smart spaces, simulation, and synthetic-data generation. Rather than functioning primarily as a general-purpose conversational model, Cosmos3-Super can use visual and action information to model possible physical-world outcomes. It is also positioned as a teacher model for distillation and post-training, where its outputs can help train smaller or more specialized systems.
NVIDIA distributes the model as open weights under the OpenMDW 1.1 license for commercial and non-commercial use, according to the supplied model information. Open weights can give research and engineering teams more control than a closed hosted service, but running a model of this size still requires considerable infrastructure.
How the model is built
Cosmos3-Super uses a Mixture-of-Transformers architecture with two complementary processing paths. An autoregressive transformer handles discrete token generation and reasoning. A diffusion transformer generates continuous modalities such as images, video, sound, and actions.
This separation supports two related workflows. The reasoner interprets multimodal inputs and can return textual reasoning, visual grounding information, temporal understanding, point localization, and bounding-box coordinates. The generator then produces simulated or predicted content, including visual scenes, audio-containing video, and action trajectories.
For a beginner, the distinction is useful: the reasoner is closer to an analysis and understanding component, while the generator is intended to create possible future states or training examples. The model is therefore better understood as a world-modeling system than as a conventional text-only large language model.
Supported inputs and outputs
The documented reasoner accepts text, text with images, and text with video. Its documented context capacity is up to 256K tokens. This long context can be useful when a task requires extended multimodal information, although the supplied research does not specify a universal maximum output size.
The generator accepts several input types:
- Text prompts
- JPG, PNG, JPEG, or WebP images
- MP4 video
- JSON action trajectories
Images and videos can use documented resolutions of 256p, 480p, and 720p. Common supported aspect ratios include 16:9, 4:3, 1:1, 3:4, and 9:16. Video inputs may include audio; NVIDIA documents stereo audio at 48 kHz for muxed MP4 inputs.
Outputs can include text, JPEG images, MP4 video, and JSON-compatible one-dimensional action lists. Generated video can contain synchronized audio encoded as stereo AAC at 48 kHz and muxed into the MP4 file. This is different from offering a separate audio-generation endpoint: the supplied documentation describes audio as part of generated video rather than as a standalone audio file.
Documented video generation supports durations from five to 400 frames, with 189 frames listed as a default generation duration. Supported frame rates include 10, 16, 24, and 30 frames per second, subject to the selected runtime and checkpoint configuration.
What Cosmos3-Super is used for
The model's main value is its ability to generate plausible physical-world futures and multimodal training data. A robotics team could condition generation on an image, video, text description, or action sequence and use the resulting scenes for simulation or post-training. An autonomous-driving team could use it to explore future visual states under documented vehicle-motion conditions.
Potentially suitable workloads include:
- Generating synthetic image and video data for perception or embodied-AI training
- Predicting or simulating future visual scenes from current observations
- Creating teacher-model outputs for distillation into smaller models
- Researching robot behavior and action-conditioned world models
- Supporting autonomous-vehicle and industrial simulation workflows
- Grounding objects and locations in images or video through points and bounding boxes
Action generation is embodiment-specific rather than universally compatible with any robot. Documented examples include camera motion, autonomous-vehicle motion, egocentric motion, single- and dual-arm robots, AgiBot platforms, UR robots, Google robots, WidowX 250, and UMI configurations. Generated actions therefore need to be mapped to a supported embodiment and validated before they are used in simulation or hardware.
Reasoning, coding, and tool support
Cosmos3-Super has meaningful multimodal reasoning capabilities, especially for visual grounding, temporal understanding, point localization, and bounding-box coordinates. Its reasoner can work with text, images, and video, making it appropriate for interpreting physical scenes and events over time.
It is not primarily a coding model. The supplied evaluation records a low coding suitability score, and the model's purpose is world modeling rather than software development. It may be useful for producing structured action information or supporting a larger Physical AI pipeline, but teams seeking code generation, debugging, or conventional programming assistance should select a model designed for those tasks.
No tool-use or function-calling capability is documented for this exact checkpoint. Action trajectories are model outputs, not evidence of general-purpose API tool use. Similarly, no verified standalone JSON mode, streaming feature, batch API, caching system, or maximum output-token limit is identified in the supplied research.
Hardware, deployment, and speed trade-offs
Cosmos3-Super is a large model intended for data-center deployment. NVIDIA lists H200, B200, and GB200 systems as recommended hardware in its model matrix. The model repository also provides multi-GPU configurations, and documented serving examples distribute the model across multiple high-memory GPUs.
It can be downloaded from NVIDIA's Hugging Face repository and deployed through NVIDIA's Cosmos software, compatible diffusion tooling, vLLM-Omni, or NVIDIA NIM-based infrastructure. These are deployment routes rather than a single guaranteed hosted API experience, so the exact setup, latency, and operating cost depend on the chosen runtime and hardware.
The principal trade-off is capability versus speed and cost. A 64-billion-parameter model that generates video, synchronized audio, and action trajectories requires substantially more compute than a small text model or an edge-focused multimodal model. The supplied evaluation rates its speed and cost suitability below its reasoning capability, reflecting this infrastructure burden. Those scores are editorial evaluations, not NVIDIA-published benchmarks.
Cosmos3-Super is consequently a better fit for organizations operating multi-GPU servers or cloud infrastructure than for users with ordinary workstations. Teams that need low-latency deployment on constrained devices should investigate smaller Cosmos models, including Cosmos3-Nano or Cosmos3-Edge, when their workload can accept reduced scale or capability.
Context limits and practical constraints
The reasoner's documented context capacity is up to 256K tokens. This is a substantial input window for multimodal reasoning, but it should not be interpreted as a promise that every workflow can process arbitrarily long or high-resolution media. Video duration, resolution, frame rate, runtime, and checkpoint configuration all affect the practical workload.
The generator supports video from five to 400 frames in the supplied documentation, but the actual computational cost rises with resolution, duration, and generation settings. The research does not identify a separate maximum output-token limit for text, so no such limit should be assumed.
Generated actions also require careful validation. A plausible trajectory is not automatically safe, physically feasible, or suitable for a particular robot. Applications should test outputs in simulation, apply task-specific constraints, and retain hardware-level safety controls. NVIDIA's default inference stack includes Cosmos Guardrail components, but guardrails do not replace application-level validation.
Pricing and availability
No official NVIDIA per-token price was identified for the Cosmos3-Super checkpoint. The model is distributed as open weights rather than presented as a conventional hosted chat model with a standard monthly plan or predictable input and output token rates.
That does not mean deployment is free. Users may incur costs for data-center GPUs, cloud instances, storage, networking, engineering, and operational support. The total cost depends on whether the model is self-hosted, deployed through an infrastructure provider, or accessed through a separate service built around NVIDIA technology. Any third-party hosting price should therefore be treated as a provider-specific service price, not as the official price of the model itself.
Main strengths and limitations
| Area | Assessment |
|---|---|
| Primary strength | Combines multimodal reasoning with world simulation, video generation, synchronized audio, and action generation. |
| Input coverage | Text, images, video, audio-containing video, and JSON action trajectories are documented. |
| Output coverage | Text, images, MP4 video with muxed audio, and embodiment-specific action lists. |
| Customization | Open-weight distribution supports self-hosting and research customization. |
| Main limitation | Requires substantial multi-GPU infrastructure and is not designed for lightweight edge use. |
| Action limitation | Action outputs are tied to documented embodiments and must be validated before deployment. |
| Pricing limitation | No official per-token price is identified for this exact checkpoint. |
When to choose Cosmos3-Super
Choose Cosmos3-Super when the project needs a large, open multimodal world model and can support the associated hardware and engineering requirements. It is particularly appropriate for Physical AI teams working on synthetic data, robotics research, autonomous-vehicle simulation, multimodal scene understanding, or teacher-model distillation.
Its open-weight format may also appeal to organizations that need to inspect, customize, or self-host a model rather than rely on a closed general-purpose API. The combination of a long-context reasoner and a multimodal generator is useful when the same research workflow needs both scene interpretation and future-state generation.
Another option is more appropriate when the priority is fast inference, low operating cost, consumer chat, general coding, edge deployment, or a simple managed API with published token pricing. NVIDIA's smaller Cosmos3-Nano and Cosmos3-Edge models may be better suited to constrained deployments, while a dedicated language or coding model is a better choice for ordinary text and software tasks.
Bottom line
Cosmos3-Super is a specialized open model for building and studying Physical AI systems. Its distinguishing capability is not ordinary conversation but the ability to connect multimodal understanding with generated visual futures, synchronized audio, and action trajectories. That makes it relevant to robotics, autonomous vehicles, simulation, and synthetic-data pipelines.
The same specialization creates its main drawbacks: high infrastructure requirements, embodiment-specific action support, no identified official per-token price, and limited relevance to conventional chat or coding. For teams with the necessary GPU resources and a physical-world modeling problem, Cosmos3-Super offers a broad open-weight foundation. For smaller, faster, or more predictable deployments, a lighter model or a purpose-built hosted service is likely to be more practical.

