MolmoAct

MolmoAct-7B-D

by Allen Institute for Artificial Intelligence (Ai2) · Available open-weight research model; superseded by MolmoAct2 as Ai2's newer MolmoAct generation

MolmoAct-7B-D is Ai2's open-weight 7B vision-language-action model for robotic manipulation. It combines image and text inputs with spatial reasoning, trajectory planning, and structured robot action prediction, while requiring task-specific integration and deployment infrastructure.

Text Actions Reasoning Coding
MolmoAct-7B-D is a specialized robotics model rather than a general chatbot or image generator. It interprets visual observations and task instructions, reasons about spatial relationships, and produces robot-oriented action predictions that can be adapted to particular hardware and manipulation tasks.
Outputs

What MolmoAct-7B-D can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
2/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MolmoAct
Model type Multimodal
Context window 4K tokens
Release date 2025-08-12
Status Available open-weight research model; superseded by MolmoAct2 as Ai2's newer MolmoAct generation
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was identified. This model is primarily a robotics action model and should not be treated as a general-purpose knowledge model.

Model notes

The canonical published checkpoint is allenai/MolmoAct-7B-D-0812. Ai2 describes MolmoAct-7B-D as its best/demo checkpoint for fine-tuning, while MolmoAct-7B-D-Pretrain-0812 is a separate pretraining checkpoint intended for inference and zero-shot evaluation. The model is based on Qwen2.5-7B with a SigLIP2 vision backbone and is designed to output structured robot actions. It is distributed as an open-weight research model rather than through an official hosted token-pricing API. The model card states that the checkpoint is released under Apache 2.0. The model was initially presented as a preview release, and Ai2 later released MolmoAct2 as a newer generation.

Model guide

MolmoAct-7B-D: An Open Vision-Language-Action Model for Robot Manipulation

MolmoAct-7B-D is an open-weight 7-billion-parameter vision-language-action model from the Allen Institute for AI. It combines camera observations and natural-language instructions with spatial reasoning, trajectory planning, and robot action prediction for manipulation tasks.

What is MolmoAct-7B-D?

MolmoAct-7B-D is an open-weight vision-language-action model developed by the Allen Institute for AI, also known as Ai2. A vision-language-action model connects visual perception and language understanding to physical behavior. In this case, the model is intended to help a robot understand what it sees, interpret an instruction, plan an interaction, and predict actions for manipulating objects.

The model is therefore aimed at robotics research and downstream adaptation, not ordinary conversation. A typical use might involve providing a camera observation and an instruction such as moving an object to a particular location. The model is designed to reason about the scene and produce structured robot-action predictions that a surrounding control system can use or further refine.

The primary published checkpoint is allenai/MolmoAct-7B-D-0812. Ai2 identifies the D variant as its best or demonstration-oriented checkpoint for downstream fine-tuning. The separate MolmoAct-7B-D-Pretrain-0812 checkpoint serves a different purpose, including inference and zero-shot evaluation, so the two should not be treated as interchangeable names for the same artifact.

Where it fits in Ai2's lineup

MolmoAct-7B-D belongs to Ai2's open-model research ecosystem and extends the Molmo direction into physical action. Unlike a general language model, it is built around the requirements of robotic manipulation: visual observations, spatial relationships, trajectories, action spaces, and adaptation to a robot platform.

It was initially presented as a preview research release. Ai2 has since released MolmoAct2 as a newer MolmoAct generation, making MolmoAct-7B-D an earlier open checkpoint rather than the newest member of the family. It can still be useful when researchers need access to the original model, its training and evaluation materials, or a checkpoint suitable for experimentation and fine-tuning.

Architecture and input limits

The model uses a Qwen2.5-7B language backbone together with a SigLIP2 vision backbone. In practical terms, the language component processes instructions and intermediate reasoning, while the vision component converts camera imagery into representations the model can use. The published configuration specifies a 378 by 378 image input size.

Its language-model context length is 4,096 tokens. This limit covers the textual and model-specific context presented during an interaction; it is not a promise that a robot can maintain an unlimited history of observations. A deployment that repeatedly sends images, instructions, intermediate states, or action information must manage the available context and the surrounding control loop.

The configuration also specifies a 256-bin action representation. This indicates that the model maps action dimensions into discrete bins rather than directly exposing an unrestricted continuous control signal. The exact interpretation of those actions depends on the model's preprocessing, action conventions, robot embodiment, and deployment code.

How its action reasoning works

MolmoAct-7B-D is designed to place intermediate spatial reasoning between perception and behavior. Instead of treating an instruction as a simple text-generation request, the model is intended to connect what is visible in the scene with where objects are located, how a trajectory might unfold, and which low-level manipulation actions are needed.

This design is important for tasks where a robot must do more than identify an object. For example, an instruction may require locating an item, understanding its relation to another object, planning a movement path, and predicting a sequence of actions that a robot arm and gripper can execute. MolmoAct-7B-D is intended to support this chain as one vision-language-action system.

The output is not limited to natural-language text. The model is designed to produce structured robot action predictions, including low-level manipulation actions. However, it should not be understood as a complete robot controller. Hardware drivers, calibration, safety checks, action execution, feedback, and task-specific adaptation remain part of the surrounding system.

Supported modalities and capabilities

CapabilityWhat is supported or known
Text inputYes; natural-language task instructions are part of the intended input.
Image inputYes; visual observations are central to the model's design.
Audio inputNot identified in the supplied model documentation.
Video inputNot identified as a native input modality for this checkpoint.
Text outputText may be produced as part of the model interface, but text is not its only intended output.
Robot action outputYes; structured action predictions are a core purpose of the model.
Image, audio, or video generationNo; it is not a generative media model.
Tool or function callingNo official tool-use or function-calling capability is identified.
Fine-tuningYes; the checkpoint is intended for downstream post-training and robot-specific fine-tuning.

The action-output capability is distinct from conventional multimodal chat. MolmoAct-7B-D does not simply describe an image or answer a question about it. Its purpose is to connect visual and language inputs with a representation of robot behavior. The model card and repository should be treated as the authority for the exact preprocessing and action-decoding workflow used with a particular checkpoint.

Training, evaluation, and deployment

Ai2 publishes the model weights, training code, datasets, and evaluation materials. The project includes workflows for simulation and real-robot settings, with reported experiments involving SimplerEnv, LIBERO, and physical robot evaluations. These are research results for particular environments and setups. They should not be read as a guarantee that the model will work without adaptation on a different robot, camera arrangement, gripper, calibration scheme, or task distribution.

Local inference uses the Transformers ecosystem with custom model code. The published checkpoint is stored in FP32, while the documentation describes casting to BF16 for inference when appropriate. The practical memory and speed requirements will depend on precision, GPU hardware, image processing, batch size, and the frequency of the robot-control loop.

Because the model is distributed as a checkpoint rather than an official hosted inference product, deployment normally requires managing the model runtime and the robot system yourself or using third-party infrastructure. There is no official per-token hosted API price listed for MolmoAct-7B-D. The cost of using it depends on local hardware, cloud GPU rental, software infrastructure, storage, and the physical robot platform. Open weights can reduce licensing and per-request API costs, but they do not make hardware, engineering, or operation free.

Main strengths and trade-offs

  • Robotics-specific design: The model is built around manipulation, spatial reasoning, trajectories, and action prediction rather than general text generation.
  • Open research materials: Weights, code, datasets, and evaluation resources are published, making it more inspectable and adaptable than a closed hosted controller.
  • Useful starting point for fine-tuning: The D checkpoint is positioned for downstream adaptation to robot-specific data and tasks.
  • Multimodal action reasoning: It combines visual observations and language instructions with a structured action representation.
  • Deployment flexibility: Researchers can run the checkpoint locally or integrate it into their own robotics stack instead of depending on a proprietary API.

These advantages come with substantial engineering trade-offs. A local open model can provide control over data, weights, and inference, but the user must handle installation, GPU resources, model integration, calibration, safety, and monitoring. A hosted general-purpose model may be easier to access, while a smaller task-specific controller may be faster and simpler for a narrowly defined robot behavior.

Limitations and safety considerations

MolmoAct-7B-D is not a general-purpose language model. Its coding capability is incidental rather than a design goal, and it should not be selected for software development, general chat, speech applications, image generation, or video generation. No official JSON mode, caching feature, streaming interface, batch API, or tool-calling interface is identified in the supplied documentation.

The model also has no published maximum output-token limit in the supplied research. Its 4,096-token context length should not be confused with a separately documented output allowance. The 256-bin action representation describes the action encoding, not a universal maximum duration or number of robot movements in every deployment.

Physical performance depends on factors outside the checkpoint: camera placement, depth or spatial inputs, calibration, robot embodiment, gripper design, action scaling, preprocessing, feedback control, and task-specific data. A prediction that looks plausible in a simulation may fail on a real robot because of occlusion, latency, friction, object variation, or an incorrect coordinate transformation.

For that reason, the model should be evaluated in simulation and controlled test environments before it is connected to physical equipment. Safety interlocks and an independent execution layer are especially important because the model should not be treated as an unattended guarantee of safe behavior.

When to choose MolmoAct-7B-D

Choose MolmoAct-7B-D when the project needs an openly published vision-language-action checkpoint for research, reproducibility, or downstream fine-tuning. It is a reasonable fit for teams studying vision-guided manipulation, spatial reasoning, trajectory planning, or action prediction and that can operate a local model and build the surrounding robotics infrastructure.

It is less suitable when the priority is a polished conversational assistant, a managed multimodal API, low-effort integration, built-in web access, speech processing, or guaranteed production uptime. For a narrowly defined production task, a smaller conventional controller or a task-specific policy may offer better speed and predictability. For newer MolmoAct research, Ai2's MolmoAct2 may be the more relevant starting point, although the supplied information does not establish that it is interchangeable with MolmoAct-7B-D or superior for every task.

In short, MolmoAct-7B-D is best viewed as an open robotics foundation checkpoint: valuable for inspection and adaptation, but incomplete as a turnkey robot product. Its central benefit is the combination of visual and language reasoning with structured action prediction; its central cost is the engineering work required to turn those predictions into reliable, safe behavior on a specific robot.


Answers to Frequently Asked Questions

What are the main technical specifications and limitations of MolmoAct-7B-D?
MolmoAct-7B-D uses a Qwen2.5-7B language backbone and a SigLIP2 vision backbone, with a published image input size of 378 by 378 pixels, a 4,096-token context length, and a 256-bin action representation. Its real-world performance depends on the robot platform, camera setup, calibration, preprocessing, action scaling, feedback control, and task-specific adaptation.
Can MolmoAct-7B-D control a robot without additional software or hardware?
No. MolmoAct-7B-D is a research checkpoint, not a complete robot controller. Deployment requires surrounding systems for preprocessing, action decoding, hardware drivers, calibration, feedback control, safety checks, and execution. The model should be tested in simulation and controlled environments before being connected to physical equipment.
What inputs and outputs does MolmoAct-7B-D support?
The model is designed to accept natural-language instructions and image-based visual observations. Its intended output includes structured robot action predictions and low-level manipulation actions, rather than only natural-language text. Audio, native video input, image generation, tool calling, and function calling are not identified as supported capabilities.
What is MolmoAct-7B-D?
MolmoAct-7B-D is an open-weight vision-language-action model developed by the Allen Institute for AI (Ai2) for robotics research. It combines visual observations and natural-language instructions to reason about scenes and predict structured robot manipulation actions.
What is the difference between allenai/MolmoAct-7B-D-0812 and MolmoAct-7B-D-Pretrain-0812?
allenai/MolmoAct-7B-D-0812 is the primary demonstration-oriented checkpoint intended for downstream fine-tuning. MolmoAct-7B-D-Pretrain-0812 serves different purposes, including inference and zero-shot evaluation, so the two checkpoints should not be treated as interchangeable.


Sources 5
Provider

About Allen Institute for Artificial Intelligence (Ai2)