Alpamayo

Alpamayo 1.5 Nano

by NVIDIA AI · Current; open weights available; also available through NVIDIA Alpamayo 1.5 NIM

NVIDIA Alpamayo 1.5 Nano is an open 10B reasoning vision-language-action model for autonomous vehicles. It combines multi-camera observations, ego-motion history and navigation text to generate future trajectories, Chain-of-Causation reasoning traces and visual question-answering responses. The model is aimed at research and development, with open weights, reference code and an NVIDIA NIM deployment option, but no published hosted price or broad context-length specification.

Text Actions Reasoning Coding
NVIDIA Alpamayo 1.5 Nano is a specialized reasoning model for autonomous-driving research and development. It is designed to connect what a vehicle sees with why a driving decision is made and what trajectory the vehicle should follow. The open-weight model is aimed at developers and researchers studying motion planning, safety evaluation, simulation and difficult road scenarios rather than users looking for a general-purpose chatbot.
Outputs

What Alpamayo 1.5 Nano can produce

Text Actions
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
2/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Alpamayo
Model type Reasoning
Maximum output 256 tokens
Release date March 16, 2026
Status Current; open weights available; also available through NVIDIA Alpamayo 1.5 NIM
Knowledge cutoff notes

NVIDIA has not published a direct knowledge-cutoff date for Alpamayo 1.5. The model is a specialized autonomous-driving VLA system trained on driving, physical-AI and reasoning data rather than a general-purpose language model with a documented cutoff.

Model notes

Canonical Hugging Face repository: nvidia/Alpamayo-1.5-10B. NVIDIA product documentation refers to the model as Alpamayo 1.5 Nano, while the repository and NGC packaging use Alpamayo-1.5-10B. The model contains an approximately 8.2B Cosmos-Reason2 backbone and a 2.3B diffusion-based action expert. It accepts multi-camera image or video inputs, ego-motion history, navigation text and questions. Its trajectory endpoint produces 64 future waypoints covering about 6.4 seconds at 10 Hz, together with reasoning text. NIM documentation lists max_tokens with a default of 256 for VLM generation; a broader model context limit is not published. The model card states that the weights are ready for non-commercial use and that commercial licensing is available upon request. Weight license: OpenMDW-1.1. Reference code license: Apache-2.0. Editorial scores reflect the model's specialized autonomous-driving role and should not be interpreted as NVIDIA benchmarks.

Model guide

NVIDIA Alpamayo 1.5 Nano: Open Reasoning for Autonomous Driving

NVIDIA Alpamayo 1.5 Nano is an open 10-billion-parameter vision-language-action model for autonomous-driving research. It combines multi-camera video, vehicle-motion history, navigation guidance and natural-language prompts to produce future driving trajectories, Chain-of-Causation reasoning traces and visual question-answering responses.

What is NVIDIA Alpamayo 1.5 Nano?

NVIDIA Alpamayo 1.5 Nano is an open vision-language-action model for autonomous vehicles. In simple terms, it can examine driving imagery, consider the vehicle's recent movement, follow navigation-related instructions and produce a proposed future driving path. It can also provide text that explains the causes behind its prediction and answer questions about the scene.

The model is part of NVIDIA's Alpamayo family of physical-AI systems. Its canonical downloadable repository is named nvidia/Alpamayo-1.5-10B, while NVIDIA product documentation refers to the model as Alpamayo 1.5 Nano. The different names describe the same current model packaging rather than two unrelated systems.

Unlike a conventional language model that mainly returns text, Alpamayo 1.5 Nano is built for a vehicle-control workflow. Its most important output is a sequence of future waypoints representing an intended motion path. The reasoning text is useful for inspection and analysis, but it is not a substitute for the rest of an autonomous-driving stack, vehicle safety controls or validation procedures.

How the model works

Alpamayo 1.5 Nano combines an 8.2-billion-parameter Cosmos-Reason2 vision-language backbone with a 2.3-billion-parameter action expert. The backbone interprets visual and language information, while the action expert uses diffusion-based trajectory decoding to generate a physically structured future path. Diffusion-based decoding is used here to produce a trajectory rather than an image.

NVIDIA describes the model as reinforcement-learning post-trained. According to the supplied model information, its training mixture includes Chain-of-Causation reasoning traces, Cosmos physical-AI data, NVIDIA autonomous-driving data and public driving data. These details are provider or model-card descriptions, not an independent assessment of real-world driving performance.

The Chain-of-Causation format is intended to make the relationship between a road situation, a driving rationale and a predicted action easier to inspect. That can be valuable when researchers are reviewing rare scenarios or investigating why a trajectory was selected. However, a readable reasoning trace should not automatically be treated as proof that the prediction is safe, correct or complete.

Inputs and outputs

The model accepts several kinds of driving context:

  • Images or video from multiple vehicle cameras
  • Ego-motion history, meaning information about the vehicle's recent movement
  • Navigation guidance or other text instructions
  • Natural-language questions about the driving scene

Alpamayo 1.5 Nano supports flexible camera counts according to the supplied documentation, allowing developers to work with different multi-camera configurations. Navigation text can provide route context or behavioral guidance, but it should be treated as advisory conditioning rather than a deterministic command to control a vehicle.

For trajectory generation, the model returns 64 future waypoints covering approximately 6.4 seconds at 10 Hz, along with vehicle-orientation information and Chain-of-Causation reasoning text. Its visual question-answering capability can answer questions about the scene, making the same model useful for analysis and evaluation in addition to path prediction.

CapabilitySupported or documented behavior
Text inputYes, including navigation guidance and questions
Image inputYes, including multi-camera observations
Video inputYes, for driving-scene observations
Audio inputNot documented
Text outputYes, including reasoning traces and visual question answers
Action outputYes, through predicted driving trajectories and waypoints
Image, audio or video generationNo

The action trajectory is a specialized machine-readable output for autonomous-driving software, not a generated image, audio clip or video. A broader context-window limit has not been published. NVIDIA's NIM documentation lists a maximum generation setting of 256 tokens by default for VLM generation, but this should not be confused with a published overall context length for the model.

Access and deployment requirements

NVIDIA provides Alpamayo 1.5 Nano as downloadable open weights through its Hugging Face organization, with reference inference code in the NVlabs Alpamayo 1.5 repository. NVIDIA also documents an Alpamayo 1.5 NIM, which exposes the model through HTTP and gRPC interfaces for trajectory prediction and visual question answering.

The downloadable model is intended for a Linux environment and requires an NVIDIA GPU with at least 24 GB of VRAM according to the supplied model card and reference repository. NVIDIA has tested the model on hardware including the H100, while the model information also lists RTX 3090, RTX 3090 Ti, RTX 4090 and A5000-class GPUs as examples of systems with enough memory. Actual performance depends on the complete software environment, input workload and hardware configuration.

This requirement makes Alpamayo 1.5 Nano fundamentally different from a hosted general-purpose AI service that can be used from a browser. Users must provide compatible compute when running the open weights themselves, or use an appropriately configured NVIDIA deployment when consuming the NIM packaging.

Main strengths and limitations

Strengths

  • Purpose-built driving outputs: The model produces future waypoints and orientation information rather than stopping at a verbal description of a scene.
  • Multi-modal driving context: It combines camera observations, motion history and language instructions in one driving-oriented workflow.
  • Inspectability: Chain-of-Causation reasoning traces can help researchers examine the relationship between a situation and a proposed action.
  • Open-weight access: Developers can download the model and reference code rather than relying only on a closed consumer interface.
  • Navigation conditioning: Route-related text can be used to study instruction-following and behavior under different driving contexts.
  • Multiple deployment paths: The model is available through open weights and a separately documented NIM interface.

Limitations

  • Specialized scope: It is not a general-purpose assistant, coding model, image generator or speech model.
  • Hardware dependence: Running the model locally requires Linux and an NVIDIA GPU with at least 24 GB of VRAM according to the supplied requirements.
  • No complete vehicle stack: It does not replace perception infrastructure, vehicle controls, safety systems, simulation, testing or operational validation.
  • Unpublished context length: NVIDIA has not supplied a broader context-window specification for the model.
  • Limited tool support: No general tool or function-calling capability is documented.
  • Commercial licensing considerations: The model card describes the weights as ready for non-commercial use, with commercial licensing available upon request.

The model's reasoning and coding scores in the supplied catalog are editorial evaluations, not NVIDIA benchmarks. The editorial assessment rates reasoning highly because reasoning traces and trajectory generation are central to the model's design, while coding suitability is low because software development is outside its intended role.

Pricing and licensing

No public per-token, subscription or recurring hosted price is specified for Alpamayo 1.5 Nano. The primary access model is downloading the open weights and running them on compatible hardware, or deploying the separately packaged Alpamayo 1.5 NIM. In that arrangement, costs come from GPU hardware, infrastructure, operations and any applicable NVIDIA licensing rather than a published consumer API rate.

The supplied model information identifies OpenMDW-1.1 as the weight license and Apache License 2.0 for the reference code. The model card states that the weights are ready for non-commercial use and that commercial licensing is available upon request. Organizations should verify the current license terms and obtain any required commercial permissions before building a production or commercial system.

Reasoning, coding and tool capabilities

Reasoning is the model's central capability, but it is specialized reasoning about driving scenes and actions. It can connect visual evidence, vehicle state, navigation guidance and a proposed trajectory, then provide Chain-of-Causation text or answer a question about the scene. This makes it more relevant to autonomous-driving analysis than a general text model, even though it is not designed to solve arbitrary reasoning problems.

Alpamayo 1.5 Nano is not intended for software development. It has no documented general-purpose coding workflow, and the supplied editorial coding score is low. It also has no documented web search, external tool use, function calling, audio processing or general agent framework. Developers can integrate its trajectory and VQA outputs into a larger system, but that integration is separate from native tool support by the model itself.

Speed and cost trade-offs

The supplied catalog gives Alpamayo 1.5 Nano a mid-range editorial speed score and a high editorial cost score. These are subjective catalog assessments, not provider-published latency or pricing benchmarks. They reflect the practical fact that a roughly 10-billion-parameter model with multi-camera inputs and trajectory reasoning requires substantial GPU resources.

Compared with a small text-only model, Alpamayo 1.5 Nano will generally demand more specialized compute and a more complex input pipeline. That additional cost is justified when the application needs driving-specific visual reasoning and trajectory output. It is difficult to justify when the task is ordinary chat, code completion, document summarization or simple classification, where a smaller or hosted model may be faster and easier to operate.

Best use cases

Alpamayo 1.5 Nano is best suited to teams working on:

  • Autonomous-driving research and development
  • Navigation-conditioned trajectory prediction
  • Motion-planning experiments and evaluation
  • Visual question answering for driving scenes
  • Safety analysis of rare or long-tail road situations
  • Simulation, data labeling and scenario investigation
  • Research into interpretable physical-AI systems

For example, a research team could provide synchronized camera observations, recent vehicle motion and a navigation instruction, then inspect both the proposed 6.4-second path and the model's explanation of the scene. Another team could ask a visual question about a road user or obstruction while comparing the answer with the predicted trajectory.

When to choose Alpamayo 1.5 Nano

Choose Alpamayo 1.5 Nano when the application needs an open, driving-specific vision-language-action model that can turn visual and motion context into future waypoints while exposing reasoning text for analysis. It is particularly attractive when local deployment, research flexibility or integration with NVIDIA GPU infrastructure matters more than a simple browser-based experience.

Choose another type of model when the priority is general conversation, programming, web research, speech, image generation or low-cost cloud inference. A conventional language model is more appropriate for text and coding tasks, while a dedicated perception, planning or control component may be preferable when a system needs narrowly validated behavior rather than a single multi-modal reasoning model.

Most importantly, Alpamayo 1.5 Nano should be treated as a development and research component. Its trajectories and explanations require testing against use-case-specific data, simulation and safety engineering before they can be considered for a real vehicle system. The model's open availability expands experimentation, but it does not remove the need for independent validation or make autonomous driving safe by itself.


Answers to Frequently Asked Questions

How can developers access and deploy Alpamayo 1.5 Nano?
NVIDIA provides downloadable open weights through its Hugging Face organization under the repository name nvidia/Alpamayo-1.5-10B, along with reference inference code in the NVlabs Alpamayo repository. NVIDIA also documents an Alpamayo 1.5 NIM with HTTP and gRPC interfaces for trajectory prediction and visual question answering.
Is Alpamayo 1.5 Nano a complete autonomous-driving system?
No. Alpamayo 1.5 Nano is a development and research component that predicts driving trajectories and explains scene-related decisions. It does not replace perception infrastructure, vehicle controls, safety systems, simulation, testing or operational validation.
What hardware is required to run NVIDIA Alpamayo 1.5 Nano?
According to the supplied model information, local deployment requires Linux and an NVIDIA GPU with at least 24 GB of VRAM. Examples include the H100, RTX 3090, RTX 3090 Ti, RTX 4090 and A5000-class GPUs, although actual performance depends on the full software and hardware configuration.
What is NVIDIA Alpamayo 1.5 Nano?
NVIDIA Alpamayo 1.5 Nano is an open vision-language-action model for autonomous driving. It analyzes camera imagery, vehicle motion history and navigation instructions to predict future driving paths, provide Chain-of-Causation reasoning text and answer questions about driving scenes.
What inputs and outputs does Alpamayo 1.5 Nano support?
The model accepts images or video from multiple cameras, ego-motion history, navigation guidance and natural-language questions. Its main action output is a trajectory containing 64 future waypoints covering approximately 6.4 seconds at 10 Hz, along with vehicle-orientation information and reasoning text.


Sources 6
Provider

About NVIDIA AI