MolmoAct2

MolmoAct2

by Allen Institute for Artificial Intelligence (Ai2) · Current open-weight foundation checkpoint

MolmoAct2 is Ai2's open vision-language-action foundation model for robot control. It combines visual and language reasoning with robot-state information and a continuous flow-matching action expert, supports fine-tuning through the LeRobot ecosystem, and includes specialized checkpoints for several robot platforms and benchmarks. The base model is intended for adaptation rather than universal plug-and-play deployment, and no hosted usage price is identified.

Text Actions Reasoning Coding
MolmoAct2 is designed for researchers and developers building robot policies rather than for general-purpose chat or a conventional cloud AI API. It combines the Molmo2-ER embodied-reasoning vision-language backbone with robot-state modeling and a continuous flow-matching action expert. The released foundation checkpoint is intended to be fine-tuned for a target robot or benchmark, while specialized checkpoints support several documented robot and simulation configurations.
Outputs

What MolmoAct2 can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
2/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MolmoAct2
Model type Multimodal
Context window 16K tokens
Release date 2026-05-05
Status Current open-weight foundation checkpoint
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was identified. The model is a robotics policy trained on embodied visual, language, state, and action data rather than a conventional general-purpose language model with a published knowledge cutoff.

Model notes

MolmoAct2 is the post-trained multi-embodiment foundation checkpoint in the MolmoAct2 family. It combines the Molmo2-ER embodied-reasoning VLM with a continuous flow-matching action expert and is intended for further fine-tuning rather than ready-to-run deployment on one specific robot. The model produces robot actions and is distributed as open weights with supporting code and datasets. Ai2 reports support for continuous and discrete action inference through its LeRobot integration. Specialized checkpoints include MolmoAct2-DROID, MolmoAct2-BimanualYAM, MolmoAct2-SO100_101, and MolmoAct2-LIBERO. Editorial scores are comparative estimates for robotics models, not provider-published ratings.

Model guide

MolmoAct2: An Open Vision-Language-Action Model for Robot Control

MolmoAct2 is an open vision-language-action foundation model from the Allen Institute for AI that combines visual reasoning, language instructions, robot-state information, and continuous action generation for adaptable robot-control policies.

What is MolmoAct2?

MolmoAct2 is an open vision-language-action model developed by the Allen Institute for AI (Ai2) for robot control. A vision-language-action model connects perception and language understanding to physical actions: it can interpret what a robot sees, process a task instruction, consider the robot's state, and produce an action sequence for the robot to execute.

That makes MolmoAct2 different from a general-purpose language model. Its output is intended to control a robot rather than simply return a written answer. The model is aimed at embodied-AI research, manipulation-policy development, simulation, and deployment experiments on supported or related hardware.

The supplied release information identifies MolmoAct2 as a current open-weight foundation checkpoint. Ai2 released it with model weights, training code, robotics datasets, evaluation rollouts, and an open action tokenizer. The model is available through its Hugging Face model repository and the accompanying official source repository.

How the model combines perception and action

MolmoAct2 accepts visual observations, a natural-language instruction, and robot-state information. Visual observations may describe the scene in front of the robot, while robot state can include information needed to understand the robot's current configuration. The model then generates machine-executable robot actions.

Its architecture combines the Molmo2-ER embodied-reasoning vision-language backbone with a continuous flow-matching action expert. In practical terms, the vision-language component helps interpret the task and scene, while the action component maps that understanding to a trajectory or set of movements. Ai2 describes the connection between these components as per-layer key-value conditioning.

The released MolmoAct2 tooling supports continuous and discrete action inference. It is integrated into Hugging Face LeRobot for training, evaluation, simulation, and deployment workflows. This integration is important because using the model generally involves more than loading a text-generation checkpoint: researchers must work with observations, robot states, action representations, datasets, simulators, or physical hardware.

Where it fits in the MolmoAct2 family

The base MolmoAct2 checkpoint is a multi-embodiment foundation model. It is not presented as a finished policy for one particular robot. Ai2 also provides specialized checkpoints for DROID Franka, bimanual YAM, SO-100/SO-101, and LIBERO environments. These variants are useful when a project matches one of those robot configurations or benchmarks more closely than the general foundation checkpoint.

Ai2 also describes MolmoAct2-Think, a related variant that adds adaptive depth-token reasoning. This is intended for tasks that benefit from more explicit three-dimensional spatial analysis. It is a positioning distinction rather than evidence that every MolmoAct2 deployment automatically performs detailed 3D reasoning.

Within Ai2's broader open-model ecosystem, MolmoAct2 is the robotics-focused member of the Molmo family. It should not be confused with a general multimodal assistant or with a model intended primarily for image description, document analysis, or conversational use.

Capabilities and supported data

CapabilityWhat the supplied specifications indicate
Visual inputSupported
Text inputSupported, including language instructions
Robot-state inputSupported as part of the action-modeling workflow
Audio inputNot identified as supported
Video inputNot verified in the model-specific data
Text outputSupported for the model's text-capable interface, where applicable
Robot-action outputSupported; this is the model's central output
Image, video, or audio generationNot supported
Tool or function callingNot identified as supported
Structured JSON outputNot identified as supported

The model-specific record lists a 16,384-token context length. No maximum output-token limit was identified. That context figure should not be interpreted as a guarantee that every robotics observation, trajectory, image representation, and software wrapper can be combined without additional preprocessing; the effective workload depends on the implementation and supported input formatting.

MolmoAct2's action output is not the same as direct image, audio, or video generation. It produces robot-control actions through the robotics stack. The model record therefore classifies multimodal output as text-only for the standardized output field, while separately identifying action output as supported.

Training and open release

The release is designed to be inspected and adapted rather than consumed only through a hosted endpoint. The supplied documentation identifies model weights, training code, robotics datasets, evaluation rollouts, and an open action tokenizer as part of the release. Training data spans multiple embodiments, camera configurations, control schemes, and task types.

Reported data sources include bimanual YAM demonstrations, DROID Franka data, SO-100/SO-101 data, BC-Z, Bridge, RT-1, and earlier MolmoAct data. This breadth helps explain why the base model is positioned as a starting point for fine-tuning across embodiments. It does not mean that the model will transfer directly to an arbitrary robot without calibration, compatible sensors, and additional training.

Fine-tuning is listed as supported. The open implementation and LeRobot integration can be useful for researchers who need to inspect training procedures, reproduce evaluations, adapt the policy, or connect it to a supported simulation or hardware setup.

Main strengths and limitations

Strengths

  • Open research access: Weights, code, datasets, evaluation material, and an action tokenizer make the system more inspectable and adaptable than a closed robot-control service.
  • Embodied reasoning: The model combines visual and language understanding with robot-state information instead of treating the task as text generation alone.
  • Multi-embodiment orientation: Training data and specialized checkpoints cover several robot configurations and benchmarks.
  • Action-generation design: The continuous flow-matching action expert is designed for robot-control outputs, with tooling for continuous and discrete action inference.
  • Research workflow integration: LeRobot support connects the model to training, evaluation, simulation, and deployment workflows.

Limitations

  • Not a ready-made universal controller: The base checkpoint is intended for fine-tuning to a target embodiment or benchmark. It should not be assumed to work immediately on an unrelated robot.
  • Action batches are not continuous re-planning: According to the supplied research, the model plans batches of actions rather than re-planning after every movement. An unexpected obstacle may therefore require another inference cycle.
  • Hardware transfer remains difficult: A substantially different robot, sensor arrangement, camera setup, or control scheme may require additional data, training, and calibration.
  • No general cloud API positioning: The release is distributed as open weights and code rather than described as a conventional hosted API with published input and output prices.
  • Safety validation is essential: A model that generates robot actions should not be treated as suitable for safety-critical autonomous operation without extensive testing, monitoring, and hardware-specific safeguards.

Reasoning, coding, and tool support

MolmoAct2 has a strong embodied-reasoning role: it must relate visual information, language instructions, robot state, and intended movement. The model record gives it an editorial reasoning score of 8 out of 10, but that score is a comparative editorial estimate, not a rating published by Ai2. The provider documentation supports the more specific claim that MolmoAct2 uses the Molmo2-ER embodied-reasoning backbone and that the MolmoAct2-Think variant adds adaptive depth-token reasoning.

Coding is not the model's primary purpose. The model record gives it an editorial coding score of 2 out of 10, reflecting that it is designed to generate robot actions rather than software. It may still be used within a coding workflow where a developer writes training, evaluation, or deployment code around it, but it should not be selected as a general programming assistant on the basis of this robotics release.

No provider-supported tool calling, function calling, web search, or general-purpose agent framework is identified. LeRobot integration provides robotics workflow support, but that is different from a model API that autonomously calls arbitrary external tools.

Pricing and access

No per-token, per-request, or recurring subscription price is identified for MolmoAct2. The model is described as an open-weight release accessible through Hugging Face and GitHub. That does not make the total cost of use zero: users may still need GPU capacity, storage, robotics hardware, simulation infrastructure, data collection, engineering work, and safety testing.

The open release can be more economical than a commercial robotics platform for teams that already operate their own infrastructure and need to fine-tune or inspect the system. Conversely, organizations without robotics engineering resources may find a managed policy service or a robot-specific commercial solution easier to operate, even if it provides less access to model internals.

When to choose MolmoAct2

Choose MolmoAct2 when you need an open robotics foundation model that can be adapted to a target embodiment, and when your team can manage training, evaluation, inference infrastructure, and physical-system validation. It is particularly relevant for:

  • robotics and embodied-AI research;
  • vision-guided manipulation;
  • fine-tuning policies for supported or closely related robot platforms;
  • experiments that require inspectable weights, code, datasets, and evaluation artifacts;
  • simulation-to-real or real-world manipulation research using the LeRobot ecosystem.

Another option may be more appropriate if you need a ready-to-run controller for a specific robot, a hosted API with predictable pricing and uptime, general-purpose coding or chat, broad tool calling, or a system that continuously re-plans around dynamic obstacles. A specialized MolmoAct2 checkpoint may also be preferable to the base model when your hardware or benchmark matches DROID Franka, bimanual YAM, SO-100/SO-101, or LIBERO more closely.

Bottom line

MolmoAct2 is best understood as an adaptable open foundation for robot-control research, not as a universal plug-and-play robot brain. Its main value comes from connecting embodied visual reasoning to action generation while providing the weights and tooling needed for fine-tuning and inspection. Its main trade-off is operational: users must handle embodiment adaptation, compute, calibration, evaluation, and safety themselves. For teams prepared to do that work, it offers a more open and customizable path than a closed robotics service; for teams seeking immediate deployment or a general AI assistant, another type of system is likely to be a better fit.


Answers to Frequently Asked Questions

Does MolmoAct2 support continuous robot control and real-time re-planning?
MolmoAct2 supports continuous and discrete action inference, but its action batches are not the same as continuous re-planning after every movement. An unexpected obstacle or scene change may require another inference cycle. Extensive hardware-specific testing, monitoring, and safety safeguards are necessary before using it for autonomous robot operation.
What does the open MolmoAct2 release include?
The release includes model weights, training code, robotics datasets, evaluation rollouts, and an open action tokenizer. It is available through the official Hugging Face model repository and GitHub source repository, with integration into Hugging Face LeRobot for training, evaluation, simulation, and deployment workflows.
What is MolmoAct2 used for?
MolmoAct2 is an open vision-language-action model for robot control. It interprets visual observations, natural-language instructions, and robot-state information to generate executable robot actions. It is designed for embodied-AI research, manipulation-policy development, simulation, fine-tuning, and deployment experiments.
Is MolmoAct2 a ready-to-use controller for any robot?
No. The base MolmoAct2 checkpoint is a multi-embodiment foundation model rather than a universal plug-and-play controller. It generally requires fine-tuning, calibration, compatible sensors, and evaluation for the target robot. Specialized checkpoints are available for DROID Franka, bimanual YAM, SO-100/SO-101, and LIBERO environments.


Sources 4
Provider

About Allen Institute for Artificial Intelligence (Ai2)