MolmoAct2

MolmoAct2-Think

by Allen Institute for Artificial Intelligence (Ai2) · Current open-weight foundation checkpoint

MolmoAct2-Think is Ai2’s open approximately 5B robotics model that predicts a compact depth representation before generating continuous robot-control actions. It is designed for depth-aware manipulation research and fine-tuning to specific robot embodiments, rather than turnkey deployment or general-purpose chat.

Text Actions Reasoning Coding
MolmoAct2-Think extends Ai2’s MolmoAct2 action-reasoning system with an intermediate depth-reasoning step. Given visual observations, language instructions, and robot-related state, it predicts a compact 10×10 discrete depth representation before generating continuous robot-control actions. This depth-aware design is aimed at embodied AI research, manipulation, and adaptation to different robot platforms.
Outputs

What MolmoAct2-Think can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
2/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MolmoAct2
Model type Robotics
Release date 2026-05-05
Status Current open-weight foundation checkpoint
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this robotics action model. Its behavior is based on trained visual-language, depth, robot-state, and action data rather than a documented general web-knowledge cutoff.

Model notes

Canonical checkpoint ID: allenai/MolmoAct2-Think. The model extends MolmoAct2 with depth-token reasoning and predicts a compact 10×10 discrete depth representation before action generation. It is a post-trained, multi-embodiment foundation checkpoint containing the visual-language model, continuous action expert, depth-token weights, and normalization metadata. Ai2 states that it is intended for further fine-tuning rather than as a ready-to-run policy for a single deployment setting. Direct ready-to-run depth-reasoning inference is provided through the fine-tuned MolmoAct2-Think-LIBERO checkpoint. The model is open-weight, approximately 5B parameters, distributed in F32 format, and released under Apache 2.0. No official hosted API pricing, context length, maximum output-token limit, or knowledge cutoff was identified. Editorial scores reflect robotics usefulness rather than general language-model performance.

Model guide

MolmoAct2-Think: Depth-Aware Vision-Language-Action Model for Robot Control

MolmoAct2-Think is an open, approximately 5-billion-parameter robotics model from Ai2 that combines visual-language reasoning, compact depth-token prediction, and continuous robot-action generation. It is intended mainly as a foundation checkpoint for fine-tuning robotic policies, not as a ready-to-deploy controller for an arbitrary robot.

What is MolmoAct2-Think?

MolmoAct2-Think is an open-weight vision-language-action model developed by the Allen Institute for AI (Ai2). Its purpose is to connect visual perception and language instructions with physical robot behavior. Instead of primarily producing conversational text, it processes a scene and task instruction, reasons over the environment, and generates actions that a robot policy can use.

The model belongs to Ai2’s MolmoAct2 family and is described as a post-trained, multi-embodiment foundation checkpoint. In practical terms, it is a starting point for researchers and developers who want to fine-tune a robotics policy for a particular robot, dataset, simulation environment, or benchmark. It is not presented as a universal controller that can be connected to any robot without additional adaptation.

Its most distinctive feature is the “Think” variant’s depth-token reasoning. Before producing robot actions, the model predicts a compact 10×10 discrete representation of depth. The resulting depth-aware information is then used by a continuous action expert, which generates robot-control actions. This creates an intermediate representation related to the relative spatial structure of the scene before action generation.

How the model works

MolmoAct2-Think combines several stages that serve different purposes:

  1. Visual and language processing: The model receives visual observations and a natural-language instruction, together with robot-related state information where required by the policy setup.
  2. Depth-token reasoning: It predicts a compact 10×10 discrete depth representation. These tokens provide a depth-aware intermediate signal rather than a full conventional depth image.
  3. Action generation: A continuous action expert uses the visual-language representation and depth-aware cache to generate robot-control actions.

The action expert is designed for closed-loop manipulation. A robot can repeatedly observe its environment and produce updated actions as the task progresses, although the exact control loop, sensors, calibration, action format, and safety mechanisms depend on the target embodiment and implementation.

The depth representation should not be confused with image generation or a general-purpose visual output. MolmoAct2-Think accepts visual input and produces machine-oriented robot actions, with depth tokens serving as an internal or intermediate reasoning signal. Its primary output is not a finished image, video, or audio file.

Where it fits in Ai2’s model lineup

MolmoAct2-Think is a specialized member of Ai2’s MolmoAct2 robotics family. Ai2’s broader open-model ecosystem includes language, multimodal, speech, document-processing, and scientific research projects, but this checkpoint is focused specifically on embodied AI and robot control.

Within the MolmoAct2 family, the Think checkpoint adds depth-token reasoning to the action-generation architecture. Ai2 also identifies MolmoAct2-Think-LIBERO as a fine-tuned checkpoint for direct depth-reasoning inference in the supported LIBERO benchmark configuration. That distinction is important: the base MolmoAct2-Think model is a general foundation checkpoint for further adaptation, while a fine-tuned variant may be the more practical starting point for a particular benchmark or robot setup.

Technical specifications and supported inputs

SpecificationAvailable information
ProviderAllen Institute for AI (Ai2)
Model familyMolmoAct2
Approximate size5 billion parameters
Model typeVision-language-action robotics model
Primary inputsVisual observations, text instructions, and robot-related state information
Primary outputsContinuous robot-control actions
Intermediate reasoningCompact 10×10 discrete depth representation
Image inputSupported
Text inputSupported
Audio or video inputNot identified for this checkpoint
Direct image, video, or audio outputNot supported as the model’s documented output role
Hosted APINo official hosted API identified
Context lengthNot published in the supplied sources
Maximum output tokensNot applicable or not published for the robotics action interface

The model card identifies a canonical checkpoint ID of allenai/MolmoAct2-Think. The checkpoint is distributed in F32 format and is approximately 5B parameters. The supplied documentation does not publish a conventional language-model context window, maximum output-token count, or knowledge-cutoff date. Those omissions are expected to matter less than robot-specific constraints such as observation size, action horizon, control frequency, hardware limits, and the format of the fine-tuning dataset.

Reasoning, action generation, and coding support

MolmoAct2-Think’s reasoning capability is specialized rather than conversational. Its depth-token stage gives the model a compact spatial signal to use before action generation. This can be useful for tasks in which the location, distance, or relative arrangement of objects affects the correct manipulation behavior.

Its documented action output is continuous robot control, not ordinary programming code. The model can therefore be useful inside a robotics policy pipeline, but it should not be evaluated as a coding assistant or general chat model. Ai2’s model record does not identify tool or function calling, web search, streaming, JSON mode, batch inference, or caching as supported product features. The supplied information also does not establish a standard function-calling interface.

Because it is an open-weight checkpoint, developers can inspect the release, run it in their own environment, and fine-tune it rather than relying on a metered provider endpoint. That flexibility comes with engineering responsibilities: users must provide compatible hardware or compute infrastructure, adapt the model to their robot, and implement operational safeguards.

Main strengths

  • Depth-aware action reasoning: The intermediate 10×10 depth representation gives the action policy an explicit spatial signal before continuous action generation.
  • Robotics specialization: The architecture is aimed at manipulation and embodied control rather than being a general assistant repurposed for robotics.
  • Open-weight access: The model weights are available through Ai2’s Hugging Face organization, and the associated MolmoAct2 code is published on GitHub.
  • Fine-tuning orientation: The foundation-checkpoint design supports adaptation to a target embodiment, dataset, simulation setting, or benchmark.
  • No hosted inference fee for the checkpoint itself: Ai2 does not publish standard per-token input or output pricing because this is distributed as an open-weight model rather than a metered hosted API product.

These are practical strengths for robotics researchers who need to inspect, modify, and evaluate a model locally. They are less relevant to users looking for an immediately available consumer assistant or a turnkey cloud endpoint.

Limitations and deployment considerations

The largest limitation is that the base checkpoint is not a ready-to-run policy for a single robot. Robot behavior depends on the robot’s embodiment, camera placement, calibration, action conventions, control interface, hardware limits, and task distribution. A model that performs well after fine-tuning in one environment should not be assumed to transfer directly to another.

The model also does not come with a published hosted API price, context window, maximum output-token limit, or general knowledge cutoff. This makes conventional language-model comparisons based on tokens, chat latency, or API cost inappropriate. Its practical cost is instead shaped by local compute, data collection, fine-tuning, evaluation, simulation, and physical hardware.

Depth reasoning does not eliminate the need for reliable sensing or safety engineering. Before physical deployment, developers should test the policy in simulation and controlled environments, enforce action limits, monitor unexpected behavior, and provide emergency-stop procedures. These precautions are especially important because generated actions can affect physical equipment and objects rather than only producing text.

Licensing and project status should also be checked before redistribution or commercial deployment. The supplied model information identifies an Apache 2.0 release and points to Ai2’s responsible-use guidance, while noting that project-specific terms and restrictions may apply.

Pricing and access

MolmoAct2-Think has no published per-token input price, output price, subscription tier, or standard hosted inference rate in the supplied sources. It is distributed as an open-weight model through Ai2’s Hugging Face organization, with implementation code available in the MolmoAct2 GitHub repository.

“No API price” does not mean that deployment is cost-free. Users may need GPU infrastructure, storage, robotics hardware, data collection, fine-tuning time, and engineering support. The financial advantage is control over the deployment and the absence of a documented per-request charge for the released checkpoint, not the elimination of all operating costs.

Best use cases

MolmoAct2-Think is a strong candidate for:

  • Research into depth-aware vision-language-action policies.
  • Robot manipulation experiments involving visual instructions and spatial relationships.
  • Fine-tuning a foundation checkpoint for a specific robot embodiment.
  • Embodied AI work in simulation and supported robotics benchmarks.
  • Projects that benefit from inspectable weights, code, and local experimentation.
  • Comparative research on intermediate visual or depth representations for robot control.

A fine-tuned checkpoint such as MolmoAct2-Think-LIBERO may be more appropriate when the goal is direct evaluation or inference in the supported LIBERO configuration. The base model is better suited to users prepared to perform adaptation and integration work.

When to choose MolmoAct2-Think

Choose MolmoAct2-Think when depth-aware embodied reasoning and open-weight customization are more important than turnkey deployment. It is particularly appropriate for a robotics team that can fine-tune a policy, control its inference environment, and evaluate behavior against a target embodiment.

Another type of option may be preferable when the priority is a hosted API, predictable per-request billing, broad tool integration, fast general-purpose responses, or conversational coding. MolmoAct2-Think is not positioned as a general language model, coding model, or consumer assistant. It also may not be the best choice when a project needs a policy that already works on a specific robot without additional fine-tuning.

Its speed and cost trade-off should likewise be understood in robotics terms. The checkpoint avoids dependence on a hosted token-priced service, but a roughly 5B-parameter open model in F32 format can require substantial local compute compared with a small, task-specific policy. Conversely, local execution and fine-tuning can provide more control and reproducibility than a remote service.

Bottom line

MolmoAct2-Think is a specialized open robotics foundation model whose defining addition is compact depth-token reasoning before continuous action generation. It is aimed at researchers and developers building or fine-tuning vision-language-action policies, not at users seeking a general chatbot or an immediately deployable universal robot controller.

Its main value is the combination of open access, Ai2’s MolmoAct2 robotics architecture, and an explicit depth-aware intermediate representation. Its main limitations are equally clear: deployment requires embodiment-specific adaptation, the model has no documented hosted API pricing or conventional token limits, and safe physical use requires substantial testing and engineering around the checkpoint.


Answers to Frequently Asked Questions

What is the difference between MolmoAct2-Think and MolmoAct2-Think-LIBERO?
MolmoAct2-Think is a general foundation checkpoint intended for further adaptation, while MolmoAct2-Think-LIBERO is a fine-tuned checkpoint designed for direct depth-reasoning inference in the supported LIBERO benchmark configuration.
Does MolmoAct2-Think have a hosted API or usage-based pricing?
No official hosted API or standard per-token pricing has been identified for MolmoAct2-Think. The model is distributed as an open-weight checkpoint through Ai2’s Hugging Face organization, but users still need to cover local compute, storage, robotics hardware, fine-tuning, and engineering costs.
How does depth-token reasoning work in MolmoAct2-Think?
Before generating actions, MolmoAct2-Think predicts a compact 10×10 discrete depth representation. This depth-aware intermediate signal helps the continuous action expert reason about the relative spatial arrangement of objects and produce manipulation actions.
Is MolmoAct2-Think ready to control any robot without fine-tuning?
No. MolmoAct2-Think is a foundation checkpoint that generally requires adaptation to the target robot’s embodiment, sensors, calibration, action format, hardware limits, and task distribution. It is not presented as a universal plug-and-play robot controller.
What is MolmoAct2-Think?
MolmoAct2-Think is an open-weight vision-language-action robotics model developed by the Allen Institute for AI (Ai2). It processes visual observations, language instructions, and robot-related state information to generate continuous robot-control actions.


Sources 5
Provider

About Allen Institute for Artificial Intelligence (Ai2)