What NVIDIA Cosmos-Transfer2.5-2B is
NVIDIA Cosmos-Transfer2.5-2B is a 2-billion-parameter world-generation model for Physical AI. Instead of answering questions or producing ordinary text, it generates video representing possible variations of a physical scene. Its inputs can combine a text description with visual control signals derived from recorded video or a simulator.
For example, a developer could provide a video or simulated scene together with depth, segmentation, or edge information and ask the model to produce a variation with different lighting, weather, objects, backgrounds, or other visible conditions. The goal is to preserve important geometry and motion while changing selected aspects of the world.
The model belongs to NVIDIA's Cosmos family and is built on Cosmos-Predict2.5. The Transfer2.5 line is focused on transforming and augmenting visual scenarios, making it substantially different from a general-purpose language model or a conventional text-to-video consumer application.
Primary purpose and position in NVIDIA's lineup
Cosmos-Transfer2.5-2B is aimed at developers building robotics, autonomous-driving, simulation, and synthetic-data systems. NVIDIA describes it as a controllable world model for Physical AI: software that needs to perceive, predict, or act in physical environments.
Within the Transfer2.5 family, NVIDIA documents general, autonomous-vehicle, and robot multiview control checkpoints. The 2B model covered here is therefore best understood as one checkpoint in a specialized video-generation family, not as a standalone chatbot or a broad AI assistant.
The model remains publicly accessible through NVIDIA's model endpoint, repository, documentation, and checkpoint distribution. However, the supplied NVIDIA research states that Transfer2.5 is no longer under active development after the June 1, 2026 launch of Cosmos 3. Cosmos 3 is the newer platform for future model, documentation, and support updates. Transfer2.5-2B may still be appropriate when an existing workflow depends on its checkpoint, code, or conditioning format.
How its visual conditioning works
Cosmos-Transfer2.5-2B uses adaptive multi-control conditioning. In practical terms, this means that different forms of guidance can be supplied at the same time, allowing the generated video to follow both a scene structure and a written instruction.
Supported control inputs
- Depth maps: provide distance and geometric information so the generated scene can retain spatial relationships.
- Edge maps: emphasize outlines, object boundaries, and the broad arrangement of structures.
- Segmentation maps: identify semantic regions or object categories in a scene.
- Blurred video: preserves broad composition and motion while leaving more visual detail for the model to generate.
- Text prompts: describe desired appearance, environment, weather, scenario changes, or other visual attributes.
These controls can come from a simulator such as NVIDIA Isaac Sim or from recorded real-world video. That supports both simulation-to-real workflows, where simulated data is made more visually varied, and real-to-real augmentation, where captured footage is transformed into additional scenarios.
What the model is used for
Robotics and sim-to-real data
Robotics teams can use the model to create visual variations of manipulation and navigation scenes while retaining the spatial relationships that matter to a robot. Possible variations include different objects, backgrounds, lighting conditions, or distractors. NVIDIA reports that data augmented with the model improved generalization in a real-robot evaluation involving such changes. That is a provider-reported result rather than an independent benchmark, so it should be treated as evidence of intended use rather than a universal performance guarantee.
Autonomous-vehicle simulation
The Cosmos Transfer family includes specialized autonomous-vehicle checkpoints for controllable multi-view traffic and driving scenes. These workflows can vary weather, lighting, traffic-light states, road environments, and actor configurations while maintaining control over scene geometry and camera views.
This makes the model relevant when a team needs many related driving scenarios but cannot collect every combination of road layout, weather, traffic, and lighting from the real world. It is primarily a data-generation and simulation component, not a complete autonomous-driving stack.
Synthetic video and world-state generation
Cosmos-Transfer2.5-2B can expand training and evaluation datasets for systems that operate in physical environments. It may be useful when real-world data is expensive, difficult to label, unsafe to collect, or missing important environmental variations.
The model's value comes from structured transformation. A team can begin with a scene whose geometry and motion are important, then request controlled visual alternatives instead of generating unrelated videos from an unconstrained prompt.
Modalities and capabilities
| Capability | Supported or documented behavior |
|---|---|
| Text input | Supported as a prompt for describing the desired scene or variation. |
| Image-like spatial controls | Depth, edge, and segmentation maps can guide generation. |
| Video input | Blurred video and video-derived control information are supported in documented workflows. |
| Video output | The model natively generates video. |
| Text output | Not a text-generation model; no conversational response capability is documented. |
| Audio input or output | Not documented for this model. |
| Tool or function calling | Not documented. |
| Structured JSON output | Not documented and not a target use case. |
Consequently, conventional language-model features such as a text context window, maximum output-token limit, reasoning mode, coding assistance, web search, and function calling do not meaningfully describe this model. The research records no conventional context length or maximum output-token value. It also rates reasoning and coding suitability as low, reflecting the model's intended role rather than a claim that it performs ordinary language reasoning.
Performance and hardware requirements
NVIDIA describes Cosmos-Transfer2.5-2B as 3.5 times smaller than Cosmos-Transfer1-7B while reporting improved prompt and physics alignment and less hallucination or error accumulation during long video generation. These are NVIDIA's claims about the model family and should not be confused with an independent evaluation.
The official model matrix lists approximately 65.4 GB of required GPU memory for the general checkpoint and a minimum of one GPU. That requirement is a major practical constraint. Although the parameter count is lower than the 7B predecessor referenced by NVIDIA, the model is still not a lightweight local video generator for ordinary laptops or low-memory graphics cards.
Published reference timings for a 720p, 16-frame-per-second, five-second video with segmentation control range from about 286 seconds on an NVIDIA B200 to about 2,326 seconds on an NVIDIA H20. These figures are reference timings, not a guaranteed service-level performance target. Actual speed depends on the checkpoint, resolution, control configuration, precision, sequence length, and implementation settings.
The trade-off is therefore clear: the model offers detailed spatial conditioning and specialized Physical AI workflows, but video generation can require considerable hardware and substantial processing time. Teams that need rapid interactive generation, low-cost experimentation, or modest local hardware may prefer a smaller or more specialized alternative, provided it offers the required control signals.
Pricing, access, and licensing
NVIDIA lists a free NIM endpoint for Cosmos-Transfer2.5-2B. The supplied research does not provide a token-based input or output price, and video-generation costs for self-hosted use depend on infrastructure, GPU time, storage, and deployment configuration.
NVIDIA also provides downloadable checkpoints together with open-source inference and post-training code. The model is released under the NVIDIA Open Model License, while the associated source code is released under Apache 2.0. The availability of post-training code does not mean that NVIDIA offers a conventional hosted fine-tuning API; it indicates that users can work with the provided code and model assets in their own environment.
Before deployment, users should check the current license text, endpoint terms, hardware requirements, and repository status. The original Transfer2.5 repository is described as publicly accessible but no longer under active development.
Important limitations
- It is not conversational: the model does not provide a general-purpose text-generation context window, chat interface, or token-based answer limit.
- It is computationally demanding: the documented general-checkpoint memory requirement is approximately 65.4 GB of GPU memory.
- Video quality is not guaranteed: generated clips may contain visual artifacts, temporal inconsistencies, or incorrect physical details.
- Controls do not eliminate errors: depth, edges, segmentation, blur, and text provide guidance, but they do not ensure perfect preservation of objects, motion, or physics.
- The platform is aging: Transfer2.5 is now a legacy line for new development compared with Cosmos 3.
- Results depend on the workflow: performance and usefulness vary with resolution, conditioning inputs, checkpoint, precision, and sequence length.
When to choose Cosmos-Transfer2.5-2B
Choose Cosmos-Transfer2.5-2B when your main problem is controllable video world generation and you can provide suitable spatial controls. It is a strong fit for robotics data augmentation, autonomous-vehicle simulation, structured video transformation, and Physical AI research where scene geometry and motion matter more than conversational interaction.
It is particularly relevant when you need to vary weather, lighting, objects, backgrounds, traffic states, or other visual conditions while retaining a recognizable scene structure. The free NIM endpoint can also provide a way to evaluate the model before committing to a self-hosted deployment, subject to the endpoint's current availability and terms.
Another option may be more appropriate when you need a chatbot, code generation, document analysis, web search, speech, embeddings, fast interactive responses, or operation on a low-memory GPU. New projects should also evaluate Cosmos 3 because NVIDIA identifies it as the current successor platform. Existing Transfer2.5 workflows, however, may still benefit from this checkpoint's documented controls and available source code.
Bottom line
NVIDIA Cosmos-Transfer2.5-2B is a specialized video world model for Physical AI, not a general AI assistant. Its main distinction is the ability to combine text with depth, edge, segmentation, blur, and video-derived controls to generate structured visual variations. That makes it useful for robotics, autonomous-vehicle simulation, and synthetic training data, but its high GPU-memory requirement, generation time, possible video artifacts, and legacy status should be part of any adoption decision.

