Cosmos-Transfer2.5

Cosmos-Transfer2.5-2B

by NVIDIA AI · Available; legacy relative to Cosmos 3; original repository under limited maintenance

NVIDIA Cosmos-Transfer2.5-2B is a 2B-parameter video world model that uses text, depth, edges, segmentation, blur, and video controls to generate physics-aware scene variations. It targets robotics, autonomous-vehicle simulation, and Physical AI data generation, but requires substantial GPU memory and is now a legacy platform relative to Cosmos 3.

Video generation Reasoning Coding
NVIDIA Cosmos-Transfer2.5-2B generates controllable video variations from text prompts and spatial signals such as depth, edges, segmentation, and blurred video. It can help robotics and autonomous-vehicle teams create additional training scenarios while preserving scene structure, motion, and physical relationships. The model is available as downloadable checkpoints and through a free NVIDIA NIM endpoint, although its general checkpoint requires substantial GPU memory and the Transfer2.5 line is now a legacy platform relative to Cosmos 3.
Outputs

What Cosmos-Transfer2.5-2B can produce

Video generation
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
3/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Cosmos-Transfer2.5
Model type Multimodal
Release date October 6, 2025
Status Available; legacy relative to Cosmos 3; original repository under limited maintenance
Knowledge cutoff notes

No conventional textual knowledge cutoff is published for this video world-generation model. Its behavior is determined by visual training data, model checkpoints, prompts, and conditioning inputs.

Model notes

Canonical model identity: Cosmos-Transfer2.5-2B. The model accepts text prompts and multiple spatial control inputs, including depth, edge, segmentation, and blur signals, and natively generates video. NVIDIA documents general, autonomous-vehicle, and robot multiview control checkpoints within the Transfer2.5 family. The general model matrix lists 65.4 GB of required GPU memory and one minimum GPU. NVIDIA reports improved prompt and physics alignment and reduced long-video error accumulation compared with Cosmos-Transfer1-7B. The repository states that Transfer2.5 is no longer under active development after Cosmos 3 launched on June 1, 2026, although the model, code, checkpoints, and NVIDIA NIM endpoint remain publicly accessible. Post-training is supported, but this should not be interpreted as a conventional hosted fine-tuning API.

Cost

Model pricing

Input No official token-based hosted price; NVIDIA lists a free NIM endpoint. Downloadable model is intended for self-hosted deployment.
Output No official token-based hosted price; video-generation costs depend on deployment infrastructure and hardware.
Model guide

NVIDIA Cosmos-Transfer2.5-2B: Controllable Video for Physical AI

NVIDIA Cosmos-Transfer2.5-2B is a 2-billion-parameter world-generation model that transforms text- and control-conditioned video inputs into physics-aware simulated video. It is designed for robotics, autonomous-vehicle simulation, synthetic-data generation, and other Physical AI workflows rather than conversational text generation.

What NVIDIA Cosmos-Transfer2.5-2B is

NVIDIA Cosmos-Transfer2.5-2B is a 2-billion-parameter world-generation model for Physical AI. Instead of answering questions or producing ordinary text, it generates video representing possible variations of a physical scene. Its inputs can combine a text description with visual control signals derived from recorded video or a simulator.

For example, a developer could provide a video or simulated scene together with depth, segmentation, or edge information and ask the model to produce a variation with different lighting, weather, objects, backgrounds, or other visible conditions. The goal is to preserve important geometry and motion while changing selected aspects of the world.

The model belongs to NVIDIA's Cosmos family and is built on Cosmos-Predict2.5. The Transfer2.5 line is focused on transforming and augmenting visual scenarios, making it substantially different from a general-purpose language model or a conventional text-to-video consumer application.

Primary purpose and position in NVIDIA's lineup

Cosmos-Transfer2.5-2B is aimed at developers building robotics, autonomous-driving, simulation, and synthetic-data systems. NVIDIA describes it as a controllable world model for Physical AI: software that needs to perceive, predict, or act in physical environments.

Within the Transfer2.5 family, NVIDIA documents general, autonomous-vehicle, and robot multiview control checkpoints. The 2B model covered here is therefore best understood as one checkpoint in a specialized video-generation family, not as a standalone chatbot or a broad AI assistant.

The model remains publicly accessible through NVIDIA's model endpoint, repository, documentation, and checkpoint distribution. However, the supplied NVIDIA research states that Transfer2.5 is no longer under active development after the June 1, 2026 launch of Cosmos 3. Cosmos 3 is the newer platform for future model, documentation, and support updates. Transfer2.5-2B may still be appropriate when an existing workflow depends on its checkpoint, code, or conditioning format.

How its visual conditioning works

Cosmos-Transfer2.5-2B uses adaptive multi-control conditioning. In practical terms, this means that different forms of guidance can be supplied at the same time, allowing the generated video to follow both a scene structure and a written instruction.

Supported control inputs

  • Depth maps: provide distance and geometric information so the generated scene can retain spatial relationships.
  • Edge maps: emphasize outlines, object boundaries, and the broad arrangement of structures.
  • Segmentation maps: identify semantic regions or object categories in a scene.
  • Blurred video: preserves broad composition and motion while leaving more visual detail for the model to generate.
  • Text prompts: describe desired appearance, environment, weather, scenario changes, or other visual attributes.

These controls can come from a simulator such as NVIDIA Isaac Sim or from recorded real-world video. That supports both simulation-to-real workflows, where simulated data is made more visually varied, and real-to-real augmentation, where captured footage is transformed into additional scenarios.

What the model is used for

Robotics and sim-to-real data

Robotics teams can use the model to create visual variations of manipulation and navigation scenes while retaining the spatial relationships that matter to a robot. Possible variations include different objects, backgrounds, lighting conditions, or distractors. NVIDIA reports that data augmented with the model improved generalization in a real-robot evaluation involving such changes. That is a provider-reported result rather than an independent benchmark, so it should be treated as evidence of intended use rather than a universal performance guarantee.

Autonomous-vehicle simulation

The Cosmos Transfer family includes specialized autonomous-vehicle checkpoints for controllable multi-view traffic and driving scenes. These workflows can vary weather, lighting, traffic-light states, road environments, and actor configurations while maintaining control over scene geometry and camera views.

This makes the model relevant when a team needs many related driving scenarios but cannot collect every combination of road layout, weather, traffic, and lighting from the real world. It is primarily a data-generation and simulation component, not a complete autonomous-driving stack.

Synthetic video and world-state generation

Cosmos-Transfer2.5-2B can expand training and evaluation datasets for systems that operate in physical environments. It may be useful when real-world data is expensive, difficult to label, unsafe to collect, or missing important environmental variations.

The model's value comes from structured transformation. A team can begin with a scene whose geometry and motion are important, then request controlled visual alternatives instead of generating unrelated videos from an unconstrained prompt.

Modalities and capabilities

CapabilitySupported or documented behavior
Text inputSupported as a prompt for describing the desired scene or variation.
Image-like spatial controlsDepth, edge, and segmentation maps can guide generation.
Video inputBlurred video and video-derived control information are supported in documented workflows.
Video outputThe model natively generates video.
Text outputNot a text-generation model; no conversational response capability is documented.
Audio input or outputNot documented for this model.
Tool or function callingNot documented.
Structured JSON outputNot documented and not a target use case.

Consequently, conventional language-model features such as a text context window, maximum output-token limit, reasoning mode, coding assistance, web search, and function calling do not meaningfully describe this model. The research records no conventional context length or maximum output-token value. It also rates reasoning and coding suitability as low, reflecting the model's intended role rather than a claim that it performs ordinary language reasoning.

Performance and hardware requirements

NVIDIA describes Cosmos-Transfer2.5-2B as 3.5 times smaller than Cosmos-Transfer1-7B while reporting improved prompt and physics alignment and less hallucination or error accumulation during long video generation. These are NVIDIA's claims about the model family and should not be confused with an independent evaluation.

The official model matrix lists approximately 65.4 GB of required GPU memory for the general checkpoint and a minimum of one GPU. That requirement is a major practical constraint. Although the parameter count is lower than the 7B predecessor referenced by NVIDIA, the model is still not a lightweight local video generator for ordinary laptops or low-memory graphics cards.

Published reference timings for a 720p, 16-frame-per-second, five-second video with segmentation control range from about 286 seconds on an NVIDIA B200 to about 2,326 seconds on an NVIDIA H20. These figures are reference timings, not a guaranteed service-level performance target. Actual speed depends on the checkpoint, resolution, control configuration, precision, sequence length, and implementation settings.

The trade-off is therefore clear: the model offers detailed spatial conditioning and specialized Physical AI workflows, but video generation can require considerable hardware and substantial processing time. Teams that need rapid interactive generation, low-cost experimentation, or modest local hardware may prefer a smaller or more specialized alternative, provided it offers the required control signals.

Pricing, access, and licensing

NVIDIA lists a free NIM endpoint for Cosmos-Transfer2.5-2B. The supplied research does not provide a token-based input or output price, and video-generation costs for self-hosted use depend on infrastructure, GPU time, storage, and deployment configuration.

NVIDIA also provides downloadable checkpoints together with open-source inference and post-training code. The model is released under the NVIDIA Open Model License, while the associated source code is released under Apache 2.0. The availability of post-training code does not mean that NVIDIA offers a conventional hosted fine-tuning API; it indicates that users can work with the provided code and model assets in their own environment.

Before deployment, users should check the current license text, endpoint terms, hardware requirements, and repository status. The original Transfer2.5 repository is described as publicly accessible but no longer under active development.

Important limitations

  • It is not conversational: the model does not provide a general-purpose text-generation context window, chat interface, or token-based answer limit.
  • It is computationally demanding: the documented general-checkpoint memory requirement is approximately 65.4 GB of GPU memory.
  • Video quality is not guaranteed: generated clips may contain visual artifacts, temporal inconsistencies, or incorrect physical details.
  • Controls do not eliminate errors: depth, edges, segmentation, blur, and text provide guidance, but they do not ensure perfect preservation of objects, motion, or physics.
  • The platform is aging: Transfer2.5 is now a legacy line for new development compared with Cosmos 3.
  • Results depend on the workflow: performance and usefulness vary with resolution, conditioning inputs, checkpoint, precision, and sequence length.

When to choose Cosmos-Transfer2.5-2B

Choose Cosmos-Transfer2.5-2B when your main problem is controllable video world generation and you can provide suitable spatial controls. It is a strong fit for robotics data augmentation, autonomous-vehicle simulation, structured video transformation, and Physical AI research where scene geometry and motion matter more than conversational interaction.

It is particularly relevant when you need to vary weather, lighting, objects, backgrounds, traffic states, or other visual conditions while retaining a recognizable scene structure. The free NIM endpoint can also provide a way to evaluate the model before committing to a self-hosted deployment, subject to the endpoint's current availability and terms.

Another option may be more appropriate when you need a chatbot, code generation, document analysis, web search, speech, embeddings, fast interactive responses, or operation on a low-memory GPU. New projects should also evaluate Cosmos 3 because NVIDIA identifies it as the current successor platform. Existing Transfer2.5 workflows, however, may still benefit from this checkpoint's documented controls and available source code.

Bottom line

NVIDIA Cosmos-Transfer2.5-2B is a specialized video world model for Physical AI, not a general AI assistant. Its main distinction is the ability to combine text with depth, edge, segmentation, blur, and video-derived controls to generate structured visual variations. That makes it useful for robotics, autonomous-vehicle simulation, and synthetic training data, but its high GPU-memory requirement, generation time, possible video artifacts, and legacy status should be part of any adoption decision.


Answers to Frequently Asked Questions

Should new projects use NVIDIA Cosmos-Transfer2.5-2B or Cosmos 3?
NVIDIA identifies Cosmos 3 as the successor platform and states that Transfer2.5 is no longer under active development after the June 1, 2026 launch of Cosmos 3. New projects should evaluate Cosmos 3, while existing workflows may still use Cosmos-Transfer2.5-2B when they depend on its checkpoint, code, or conditioning format.
Is NVIDIA Cosmos-Transfer2.5-2B a chatbot or text-generation model?
No. NVIDIA Cosmos-Transfer2.5-2B generates controllable video rather than conversational text. It does not document support for ordinary language-model features such as chat, text completion, function calling, structured JSON output, or coding assistance.
How much GPU memory does NVIDIA Cosmos-Transfer2.5-2B require?
NVIDIA's model matrix lists approximately 65.4 GB of GPU memory for the general Cosmos-Transfer2.5-2B checkpoint, with a minimum of one GPU. Actual requirements can vary based on resolution, precision, sequence length, checkpoint, and control configuration.
What is NVIDIA Cosmos-Transfer2.5-2B used for?
NVIDIA Cosmos-Transfer2.5-2B is a specialized video world-generation model for Physical AI. It is designed for robotics, autonomous-vehicle simulation, synthetic-data generation, and controlled visual transformation of physical scenes.
What types of inputs can NVIDIA Cosmos-Transfer2.5-2B use?
The model can combine text prompts with depth maps, edge maps, segmentation maps, blurred video, and other video-derived control signals. These inputs guide the generated video while helping preserve scene geometry, object relationships, and motion.


Sources 5
Provider

About NVIDIA AI