Isaac GR00T N1

NVIDIA Isaac GR00T N1.7

by NVIDIA AI · Current; general availability

NVIDIA Isaac GR00T N1.7 is an open vision-language-action foundation model for humanoid robot manipulation. It processes language, images, video observations, and robot state to predict executable action sequences, with a 40-step base horizon, cross-embodiment design, fine-tuning support, and ONNX or TensorRT deployment options.

Actions Reasoning Coding
NVIDIA Isaac GR00T N1.7 is designed for robots rather than conversational text generation. The model receives task instructions, camera images or video, and embodiment-specific robot state, then predicts a sequence of control actions for manipulation. Its N1.7 release uses the Cosmos-Reason2-2B vision-language backbone, supports relative end-effector actions, was pretrained with human video and robot demonstrations, and provides a 40-step base action horizon.
Outputs

What NVIDIA Isaac GR00T N1.7 can produce

Actions
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Fine-tuning Batch API Multimodal output
Model profile

Performance characteristics

4/10 Reasoning
1/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Isaac GR00T N1
Model type Multimodal
Release date 2026-07-07
Status Current; general availability
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was published in the reviewed NVIDIA repository, model card, or developer documentation. This is an action-policy model rather than a conventional standalone text-generation model.

Model notes

The exact canonical base checkpoint is nvidia/GR00T-N1.7-3B. NVIDIA also publishes fine-tuned N1.7 checkpoints for LIBERO, DROID, SimplerEnv Bridge, and SimplerEnv Fractal. N1.7 initially released as Early Access on April 18, 2026 and reached General Availability on July 7, 2026. The model is approximately 3B parameters and uses the gated nvidia/Cosmos-Reason2-2B backbone through Hugging Face. The base checkpoint has a 40-step action horizon, 132 state/action dimensions, and a flow-matching DiT action head. Model weights use the NVIDIA Open Model License; the reference code is Apache 2.0. Pricing is not applicable to the downloadable model itself.

Model guide

NVIDIA Isaac GR00T N1.7: An Open Action Model for Humanoid Robot Manipulation

NVIDIA Isaac GR00T N1.7 is an open 3-billion-parameter vision-language-action model for generalized humanoid robot manipulation. It combines language instructions, visual observations, and robot state information to predict executable action sequences across supported robot embodiments.

What Is NVIDIA Isaac GR00T N1.7?

NVIDIA Isaac GR00T N1.7 is an open vision-language-action model for generalized humanoid robotics. A vision-language-action model connects perception and language understanding to physical control: instead of ending with a text response, it predicts what a robot should do next.

In practical terms, a robot application can provide an instruction such as placing an object in a container, camera observations of the workspace, and the robot's current state. GR00T N1.7 then produces a sequence of future actions. Depending on the robot embodiment, those actions can describe joint positions, end-effector movement, gripper commands, or other robot-specific control streams.

The canonical base checkpoint is nvidia/GR00T-N1.7-3B, which contains approximately 3 billion parameters. NVIDIA publishes the model weights, reference implementation, inference utilities, fine-tuning workflows, evaluation tools, and deployment options through the Isaac-GR00T project.

Where It Fits in NVIDIA's Current AI Lineup

GR00T N1.7 is a specialized model within NVIDIA's robotics and physical-AI work, not a general-purpose chatbot or a conventional text-generation API model. It is intended to be integrated into a robot-control system, where perception, policy inference, calibration, safety monitoring, and low-level execution work together.

The model is distributed as a downloadable checkpoint and open-source implementation rather than as a primary consumer subscription product. NVIDIA also provides related N1.7 fine-tuned checkpoints for environments and datasets including LIBERO, DROID, SimplerEnv Bridge, and SimplerEnv Fractal. These are derivatives or task-specific versions of the same model family, not separate general-purpose assistants.

N1.7 was initially released as Early Access on April 18, 2026, and reached general availability on July 7, 2026, according to the supplied release information. The model uses NVIDIA Cosmos-Reason2-2B as its vision-language backbone. Access authorization is required for that backbone on Hugging Face, even though the GR00T project itself provides the relevant model and software materials.

Inputs, Outputs, and Action Horizon

GR00T N1.7 accepts three broad categories of input:

  • Language: a natural-language instruction describing the task.
  • Visual observations: camera images or video sequences from the robot's environment.
  • Robot state: embodiment-specific information such as joint, end-effector, or other control state.

Each supported embodiment defines the expected state, action, video, and language fields. This matters because a policy trained for one robot cannot automatically be treated as a drop-in controller for every other robot. The input and output configuration must match the robot's embodiment tag and interface.

The primary output is an array or tensor of future robot actions. The base N1.7 policy predicts up to 40 future action steps, an increase from the earlier 16-step horizon described for the model interface. The supplied documentation also identifies 132 state and action dimensions for the N1.7 interface. The output is therefore a structured control sequence, not text, an image, audio, or video generated for human consumption.

No conventional token context window or maximum text-output-token limit is published in the supplied sources. The 40-step action horizon should not be confused with a language-model context length: it describes the number of future robot-control steps predicted by the policy.

What Changed in the N1.7 Release?

N1.7 replaces the previous Eagle-based visual-language backbone with NVIDIA Cosmos-Reason2-2B, which uses a Qwen3-VL architecture. The supplied research identifies flexible image resolution and preservation of native image aspect ratios without fixed padding as important properties of the new backbone.

The release also introduces a relative end-effector action space. Rather than representing every movement only as an absolute target position, the policy can express changes relative to the robot's current end-effector pose. This is intended to improve transfer between robot embodiments and make it more practical to combine robot demonstrations with human-video data.

NVIDIA reports that N1.7 was pretrained with 20,000 hours of EgoScale human video alongside diverse robot demonstrations. Human video does not directly provide all the control information available from a robot demonstration, so the value of this data depends on how effectively the training and action representation connect visual behavior to executable robot movement. It should be understood as a pretraining input, not as a guarantee of reliable control in every environment.

Main Strengths and Practical Capabilities

  • Robot-oriented output: the model directly predicts action sequences suitable for a robot policy pipeline rather than requiring a separate text-to-control interpretation stage.
  • Cross-embodiment focus: relative end-effector actions and embodiment-specific interfaces are designed to support transfer across different robots.
  • Multimodal observation: language, images, video, and robot state can be combined in one policy workflow.
  • Human-video pretraining: the N1.7 training recipe incorporates a large reported volume of human video in addition to robot demonstrations.
  • Fine-tuning support: developers can adapt the model to custom tasks, robots, datasets, and environments using robot demonstrations and LeRobot-compatible datasets.
  • Deployment flexibility: NVIDIA documents local inference, a remote policy-server arrangement, PyTorch execution, ONNX export, and TensorRT acceleration.
  • Batch inference: the policy tools support processing multiple inference examples, which can be useful during evaluation and training workflows.

These strengths make GR00T N1.7 most relevant to robotics researchers, robot manufacturers, integrators, and developers building manipulation systems. They are less relevant to someone looking for a ready-to-use conversational assistant.

Fine-Tuning and Deployment Requirements

Fine-tuning allows a developer to adapt the pretrained policy using demonstrations from a particular robot or task. Multiple dataset paths and mixture weights can be configured for combined training. NVIDIA's documented minimum fine-tuning configuration requires a GPU with at least 40 GB of VRAM, so adapting the model is not a lightweight laptop workflow.

For deployment, a robot-side client can communicate with a remote policy server. This separates high-compute neural inference from the part of the system that interfaces directly with the robot. Such an arrangement can simplify hardware placement, but network reliability and latency become part of the control design.

ONNX and TensorRT export provide options for optimized inference. TensorRT is particularly relevant when the policy must respond quickly in a closed-loop system, where the robot repeatedly observes the scene, predicts actions, executes them, and observes the result again. Exact throughput and latency vary with hardware, precision, observation configuration, and deployment settings; the supplied research does not establish one universal performance number.

Reasoning, Coding, and Tool Support

GR00T N1.7 uses a vision-language backbone capable of processing visual and language information, but it is not presented as a general-purpose reasoning model with a published reasoning benchmark or user-facing chain-of-thought feature. Its useful form of reasoning is task-conditioned policy inference: connecting observations and instructions to a likely sequence of robot actions.

Coding is not a target capability. Developers use Python-based project tools and model code to train or deploy the policy, but the model itself is not intended to generate software in the way a coding assistant does. Similarly, the supplied specifications do not identify general tool calling, web search, function calling, or plug-in support. The robot-control interface is the model's primary integration mechanism.

Pricing and Cost Trade-offs

NVIDIA does not advertise a recurring per-token or per-request price for the downloadable GR00T N1.7 checkpoint in the supplied research. Pricing is therefore not applicable to the model weights themselves in the same way it would be for a hosted language-model API. The model weights use the NVIDIA Open Model License, while the reference code uses Apache 2.0.

Free access to the weights does not mean that deployment has no cost. Users may need a GPU with sufficient memory, storage for model files, compatible software, and additional robot hardware. Fine-tuning has a documented minimum of 40 GB of GPU VRAM. Inference can be run locally or through a server, and TensorRT may improve speed on supported NVIDIA hardware, but the actual infrastructure cost depends on the deployment design.

Compared with a small conventional policy model, GR00T N1.7 may require more compute because it combines a multimodal backbone with action prediction. Compared with a hosted robotics service, local deployment can offer more control over data and latency but shifts hardware, maintenance, and safety responsibilities to the implementer. The supplied sources do not provide a standardized cost-per-action or speed comparison.

Limitations and Safety Considerations

GR00T N1.7 is not a guarantee of autonomous robot competence. Results depend on the robot embodiment, camera placement, image quality, calibration, training data, action representation, inference latency, and the physical environment.

NVIDIA documents weaker performance in situations including cluttered scenes, long-horizon packing, precise shelf placement, and sequences that require recovery after a failed grasp. These limitations are important because a policy that performs well on familiar demonstrations may still fail when objects are arranged differently, visibility is blocked, or an earlier action produces an unexpected state.

The model is not intended for mission-critical functional-safety applications without additional safeguards. A real deployment should add collision prevention, workspace limits, emergency stops, action validation, monitoring, recovery logic, and robot-specific safety controls. The policy output should be treated as one component of a controlled robotics system, not as an independently trusted safety controller.

When to Choose NVIDIA Isaac GR00T N1.7

Choose GR00T N1.7 when the main problem is generalized humanoid robot manipulation and you need an adaptable vision-language-action policy that can be fine-tuned or deployed on your own infrastructure. It is a particularly suitable candidate for research involving cross-embodiment transfer, robot demonstrations, human-video pretraining, multimodal observation, and action-sequence prediction.

It may also be appropriate when TensorRT or ONNX deployment, local inference, and a remote policy-server architecture are useful requirements. The open checkpoint and reference code provide more control than a closed hosted service, although they also require more engineering work.

Another option may be more appropriate if the goal is general conversation, software coding, image generation, audio processing, web research, or a simple consumer-facing assistant. A smaller or more specialized robotics policy may also be preferable when hardware is limited, the task is narrow, or predictable low-latency behavior matters more than broad multimodal generalization. For safety-critical operation, no model should replace an independently designed safety system.


Answers to Frequently Asked Questions

Is NVIDIA Isaac GR00T N1.7 safe for autonomous or mission-critical robot operation?
GR00T N1.7 should not be treated as an independently trusted safety controller or as a guarantee of autonomous robot competence. Real deployments should add collision prevention, workspace limits, emergency stops, action validation, monitoring, recovery logic, and robot-specific safety controls, especially because performance can decline in cluttered scenes, long-horizon tasks, precise placement, and failed-grasp recovery.
What hardware and deployment options are required for GR00T N1.7?
Fine-tuning requires a GPU with at least 40 GB of VRAM according to NVIDIA's documented minimum configuration. Deployment can use local inference, a remote policy server, PyTorch, ONNX export, or TensorRT acceleration. Actual requirements and performance depend on the robot, hardware, observation setup, precision, and latency targets.
What changed in NVIDIA Isaac GR00T N1.7 compared with earlier versions?
N1.7 replaces the earlier Eagle-based vision-language backbone with NVIDIA Cosmos-Reason2-2B, which uses a Qwen3-VL architecture. It also introduces relative end-effector actions, supports flexible image resolution and native aspect ratios, and was pretrained using 20,000 hours of EgoScale human video alongside robot demonstrations.
What is NVIDIA Isaac GR00T N1.7 used for?
NVIDIA Isaac GR00T N1.7 is an open vision-language-action model for generalized humanoid robot manipulation. It combines natural-language instructions, camera observations, and robot state to predict future actions such as joint movements, end-effector commands, and gripper actions.
What are the inputs and outputs of GR00T N1.7?
GR00T N1.7 accepts a language instruction, visual observations such as images or video, and embodiment-specific robot state. Its primary output is a structured sequence of up to 40 future robot actions, with the N1.7 interface documenting 132 state and action dimensions.


Sources 6
Provider

About NVIDIA AI