What is NVIDIA Alpamayo 1.5 Nano?
NVIDIA Alpamayo 1.5 Nano is an open vision-language-action model for autonomous vehicles. In simple terms, it can examine driving imagery, consider the vehicle's recent movement, follow navigation-related instructions and produce a proposed future driving path. It can also provide text that explains the causes behind its prediction and answer questions about the scene.
The model is part of NVIDIA's Alpamayo family of physical-AI systems. Its canonical downloadable repository is named nvidia/Alpamayo-1.5-10B, while NVIDIA product documentation refers to the model as Alpamayo 1.5 Nano. The different names describe the same current model packaging rather than two unrelated systems.
Unlike a conventional language model that mainly returns text, Alpamayo 1.5 Nano is built for a vehicle-control workflow. Its most important output is a sequence of future waypoints representing an intended motion path. The reasoning text is useful for inspection and analysis, but it is not a substitute for the rest of an autonomous-driving stack, vehicle safety controls or validation procedures.
How the model works
Alpamayo 1.5 Nano combines an 8.2-billion-parameter Cosmos-Reason2 vision-language backbone with a 2.3-billion-parameter action expert. The backbone interprets visual and language information, while the action expert uses diffusion-based trajectory decoding to generate a physically structured future path. Diffusion-based decoding is used here to produce a trajectory rather than an image.
NVIDIA describes the model as reinforcement-learning post-trained. According to the supplied model information, its training mixture includes Chain-of-Causation reasoning traces, Cosmos physical-AI data, NVIDIA autonomous-driving data and public driving data. These details are provider or model-card descriptions, not an independent assessment of real-world driving performance.
The Chain-of-Causation format is intended to make the relationship between a road situation, a driving rationale and a predicted action easier to inspect. That can be valuable when researchers are reviewing rare scenarios or investigating why a trajectory was selected. However, a readable reasoning trace should not automatically be treated as proof that the prediction is safe, correct or complete.
Inputs and outputs
The model accepts several kinds of driving context:
- Images or video from multiple vehicle cameras
- Ego-motion history, meaning information about the vehicle's recent movement
- Navigation guidance or other text instructions
- Natural-language questions about the driving scene
Alpamayo 1.5 Nano supports flexible camera counts according to the supplied documentation, allowing developers to work with different multi-camera configurations. Navigation text can provide route context or behavioral guidance, but it should be treated as advisory conditioning rather than a deterministic command to control a vehicle.
For trajectory generation, the model returns 64 future waypoints covering approximately 6.4 seconds at 10 Hz, along with vehicle-orientation information and Chain-of-Causation reasoning text. Its visual question-answering capability can answer questions about the scene, making the same model useful for analysis and evaluation in addition to path prediction.
| Capability | Supported or documented behavior |
|---|---|
| Text input | Yes, including navigation guidance and questions |
| Image input | Yes, including multi-camera observations |
| Video input | Yes, for driving-scene observations |
| Audio input | Not documented |
| Text output | Yes, including reasoning traces and visual question answers |
| Action output | Yes, through predicted driving trajectories and waypoints |
| Image, audio or video generation | No |
The action trajectory is a specialized machine-readable output for autonomous-driving software, not a generated image, audio clip or video. A broader context-window limit has not been published. NVIDIA's NIM documentation lists a maximum generation setting of 256 tokens by default for VLM generation, but this should not be confused with a published overall context length for the model.
Access and deployment requirements
NVIDIA provides Alpamayo 1.5 Nano as downloadable open weights through its Hugging Face organization, with reference inference code in the NVlabs Alpamayo 1.5 repository. NVIDIA also documents an Alpamayo 1.5 NIM, which exposes the model through HTTP and gRPC interfaces for trajectory prediction and visual question answering.
The downloadable model is intended for a Linux environment and requires an NVIDIA GPU with at least 24 GB of VRAM according to the supplied model card and reference repository. NVIDIA has tested the model on hardware including the H100, while the model information also lists RTX 3090, RTX 3090 Ti, RTX 4090 and A5000-class GPUs as examples of systems with enough memory. Actual performance depends on the complete software environment, input workload and hardware configuration.
This requirement makes Alpamayo 1.5 Nano fundamentally different from a hosted general-purpose AI service that can be used from a browser. Users must provide compatible compute when running the open weights themselves, or use an appropriately configured NVIDIA deployment when consuming the NIM packaging.
Main strengths and limitations
Strengths
- Purpose-built driving outputs: The model produces future waypoints and orientation information rather than stopping at a verbal description of a scene.
- Multi-modal driving context: It combines camera observations, motion history and language instructions in one driving-oriented workflow.
- Inspectability: Chain-of-Causation reasoning traces can help researchers examine the relationship between a situation and a proposed action.
- Open-weight access: Developers can download the model and reference code rather than relying only on a closed consumer interface.
- Navigation conditioning: Route-related text can be used to study instruction-following and behavior under different driving contexts.
- Multiple deployment paths: The model is available through open weights and a separately documented NIM interface.
Limitations
- Specialized scope: It is not a general-purpose assistant, coding model, image generator or speech model.
- Hardware dependence: Running the model locally requires Linux and an NVIDIA GPU with at least 24 GB of VRAM according to the supplied requirements.
- No complete vehicle stack: It does not replace perception infrastructure, vehicle controls, safety systems, simulation, testing or operational validation.
- Unpublished context length: NVIDIA has not supplied a broader context-window specification for the model.
- Limited tool support: No general tool or function-calling capability is documented.
- Commercial licensing considerations: The model card describes the weights as ready for non-commercial use, with commercial licensing available upon request.
The model's reasoning and coding scores in the supplied catalog are editorial evaluations, not NVIDIA benchmarks. The editorial assessment rates reasoning highly because reasoning traces and trajectory generation are central to the model's design, while coding suitability is low because software development is outside its intended role.
Pricing and licensing
No public per-token, subscription or recurring hosted price is specified for Alpamayo 1.5 Nano. The primary access model is downloading the open weights and running them on compatible hardware, or deploying the separately packaged Alpamayo 1.5 NIM. In that arrangement, costs come from GPU hardware, infrastructure, operations and any applicable NVIDIA licensing rather than a published consumer API rate.
The supplied model information identifies OpenMDW-1.1 as the weight license and Apache License 2.0 for the reference code. The model card states that the weights are ready for non-commercial use and that commercial licensing is available upon request. Organizations should verify the current license terms and obtain any required commercial permissions before building a production or commercial system.
Reasoning, coding and tool capabilities
Reasoning is the model's central capability, but it is specialized reasoning about driving scenes and actions. It can connect visual evidence, vehicle state, navigation guidance and a proposed trajectory, then provide Chain-of-Causation text or answer a question about the scene. This makes it more relevant to autonomous-driving analysis than a general text model, even though it is not designed to solve arbitrary reasoning problems.
Alpamayo 1.5 Nano is not intended for software development. It has no documented general-purpose coding workflow, and the supplied editorial coding score is low. It also has no documented web search, external tool use, function calling, audio processing or general agent framework. Developers can integrate its trajectory and VQA outputs into a larger system, but that integration is separate from native tool support by the model itself.
Speed and cost trade-offs
The supplied catalog gives Alpamayo 1.5 Nano a mid-range editorial speed score and a high editorial cost score. These are subjective catalog assessments, not provider-published latency or pricing benchmarks. They reflect the practical fact that a roughly 10-billion-parameter model with multi-camera inputs and trajectory reasoning requires substantial GPU resources.
Compared with a small text-only model, Alpamayo 1.5 Nano will generally demand more specialized compute and a more complex input pipeline. That additional cost is justified when the application needs driving-specific visual reasoning and trajectory output. It is difficult to justify when the task is ordinary chat, code completion, document summarization or simple classification, where a smaller or hosted model may be faster and easier to operate.
Best use cases
Alpamayo 1.5 Nano is best suited to teams working on:
- Autonomous-driving research and development
- Navigation-conditioned trajectory prediction
- Motion-planning experiments and evaluation
- Visual question answering for driving scenes
- Safety analysis of rare or long-tail road situations
- Simulation, data labeling and scenario investigation
- Research into interpretable physical-AI systems
For example, a research team could provide synchronized camera observations, recent vehicle motion and a navigation instruction, then inspect both the proposed 6.4-second path and the model's explanation of the scene. Another team could ask a visual question about a road user or obstruction while comparing the answer with the predicted trajectory.
When to choose Alpamayo 1.5 Nano
Choose Alpamayo 1.5 Nano when the application needs an open, driving-specific vision-language-action model that can turn visual and motion context into future waypoints while exposing reasoning text for analysis. It is particularly attractive when local deployment, research flexibility or integration with NVIDIA GPU infrastructure matters more than a simple browser-based experience.
Choose another type of model when the priority is general conversation, programming, web research, speech, image generation or low-cost cloud inference. A conventional language model is more appropriate for text and coding tasks, while a dedicated perception, planning or control component may be preferable when a system needs narrowly validated behavior rather than a single multi-modal reasoning model.
Most importantly, Alpamayo 1.5 Nano should be treated as a development and research component. Its trajectories and explanations require testing against use-case-specific data, simulation and safety engineering before they can be considered for a real vehicle system. The model's open availability expands experimentation, but it does not remove the need for independent validation or make autonomous driving safe by itself.

