MolmoAct

MolmoAct 7B-D Pretrain

by Allen Institute for Artificial Intelligence (Ai2) · Available open-weight preview checkpoint

An open-weight AllenAI robotics checkpoint based on Qwen2.5-7B and SigLIP2. It accepts image observations and language instructions, generates spatial reasoning traces, and produces parsed actions for compatible robot manipulation environments. The model is intended for research, SimplerEnv evaluation, and downstream adaptation rather than general chat or turnkey hardware control.

Text Actions Reasoning Coding
MolmoAct 7B-D Pretrain is a research-focused vision-language-action model for robotic manipulation. It accepts an instruction and visual observations, produces a textual reasoning trace involving depth and trajectories, and exposes representations that can be parsed into actions for compatible robot environments. The checkpoint is openly available for researchers who want to reproduce experiments, adapt the model to another robot, or continue training it on task-specific data.
Outputs

What MolmoAct 7B-D Pretrain can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
2/10 Coding
4/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MolmoAct
Model type Multimodal
Maximum output 256 tokens
Release date 2025-08-12
Status Available open-weight preview checkpoint
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this robotics checkpoint. Its training data and release date should not be treated as a knowledge-cutoff specification.

Model notes

Canonical Hugging Face checkpoint ID is allenai/MolmoAct-7B-D-Pretrain-0812. The model is based on Qwen2.5-7B with a SigLIP2 vision backbone and is pre-trained on the MolmoAct pre-training mixture. It generates textual reasoning traces, depth tokens, trajectory traces, and representations that can be parsed into robot actions; action output is not ordinary tool calling. The checkpoint is distributed in FP32 and can be cast to BF16 for inference. It is intended for downstream mid-training or reproducing zero-shot SimplerEnv results. The model card identifies it as a preview release and licenses it under Apache 2.0 for research and educational use. The 256-token value is the published quick-start generation example rather than a verified architectural maximum output limit.

Cost

Model pricing

Input No official hosted API pricing; self-hosted checkpoint
Output No official hosted API pricing; self-hosted checkpoint
Model guide

MolmoAct 7B-D Pretrain: Open Vision-Language-Action Reasoning for Robot Manipulation

MolmoAct 7B-D Pretrain is an open-weight robotics checkpoint from the Allen Institute for AI that combines image understanding, language instructions, spatial reasoning, and robot action generation. Built on Qwen2.5-7B with a SigLIP2 vision backbone, it is intended for downstream mid-training and reproducing zero-shot SimplerEnv manipulation results rather than general chat or hosted API use.

What MolmoAct 7B-D Pretrain is

MolmoAct 7B-D Pretrain is an open-weight robotics model provided by the Allen Institute for AI (Ai2). Its purpose is narrower than that of a general-purpose conversational model: it is designed to connect visual observations and natural-language instructions to actions for robot manipulation.

In a typical use case, a camera observes a tabletop or home environment while a user or program supplies an instruction such as moving an object. The model interprets the scene, reasons about the required movement, and generates output that includes a textual trace, depth-related information, trajectory information, and an action representation that can be parsed by compatible robotics code.

The canonical published checkpoint is allenai/MolmoAct-7B-D-Pretrain-0812 on Hugging Face. The model card describes it as an inference checkpoint for reproducing zero-shot results in SimplerEnv, including the Google Robot environment, and as a base checkpoint for downstream mid-training. It is therefore best understood as a research and adaptation starting point, not as a finished controller for arbitrary physical robots.

Where it fits in Ai2's catalog

MolmoAct belongs to Ai2's open-model research work and extends the Molmo direction into vision-language-action robotics. Its role is distinct from a general language model, a visual question-answering model, or a consumer assistant. The checkpoint is meant to be loaded and run with research code, typically alongside a compatible simulated or physical robotics environment.

The related MolmoAct family includes checkpoints intended for different stages of the robotics workflow. This pre-training checkpoint is particularly relevant when researchers want to continue training or reproduce the published pre-training-to-evaluation pipeline. The supplied research indicates that users adapting the system to a real robot may prefer a task-specialized MolmoAct checkpoint when available, rather than treating the pre-training release as a ready-made deployment model.

Architecture and training details

MolmoAct 7B-D Pretrain is based on Qwen2.5-7B, a seven-billion-parameter language-model foundation, and uses SigLIP2 as its vision backbone. The vision component is initialized using the Molmo pre-training approach. In practical terms, the language component handles instruction interpretation and structured reasoning, while the vision component converts camera images into information the model can use when planning an action.

The model was pre-trained on the MolmoAct pre-training mixture. According to the supplied research, that mixture combines action-reasoning data derived from robotics datasets with auxiliary robot and web data. The broader MolmoAct data work includes approximately 10,000 high-quality trajectories from a single-arm Franka robot carrying out 93 manipulation tasks in home and tabletop environments. These details help explain the model's intended operating range, but they should not be read as proof that it will work without adaptation on every robot, camera arrangement, or task.

The published checkpoint is distributed in FP32 format. The official inference example can cast the weights to BF16 at load time, which may reduce memory requirements and improve practical inference performance on hardware that supports BF16. The supplied material does not verify a fixed architectural context window or a maximum generation limit.

Inputs, outputs, and action reasoning

The verified input pattern is multimodal: the model accepts text instructions together with image observations. The published workflow uses the Transformers library and an image-text-to-text model interface. It asks the model to reason about elements such as depth, the end-effector trajectory, and the action needed to complete a manipulation instruction.

The output is not simply a chat response. MolmoAct generates a textual reasoning trace as well as depth perception tokens and trajectory information. Custom parsing methods then convert the generated action representation into robot actions. The research includes a seven-value action example containing positional, rotational, and gripper-related values. The exact interpretation of those values depends on the compatible robotics setup and its normalization and control conventions.

This is an important distinction for developers: the action representation is not ordinary function calling or generic tool use. The model does not provide a verified general-purpose tool API in the supplied specifications. Instead, it produces robotics-oriented output that the MolmoAct code and an appropriate environment can interpret.

Verified capabilities and unknowns

AreaWhat the supplied research supports
Text inputSupported for natural-language manipulation instructions.
Image inputSupported for visual observations used in action reasoning.
Text outputSupported, including a generated reasoning trace.
Robot action outputSupported through parsed action representations for compatible environments.
Audio and video inputNot verified for this checkpoint.
Image, audio, or video generationNot supported according to the supplied model record.
Tool or function callingNot verified; robotics action parsing is a different mechanism.
Structured JSON outputNot verified as a distinct model capability.
Context lengthNo authoritative context limit was supplied.
Maximum outputNo verified architectural maximum was supplied. A 256-token value appears in the quick-start generation example and should not be treated as a confirmed maximum.

The model's reasoning capability is best described in operational terms: it generates spatially relevant traces involving depth and trajectories before producing an action representation. That makes its intermediate output inspectable, but a reasoning trace should not be treated as a guarantee that the planned action is correct or safe.

Main strengths

  • Purpose-built robotics focus: The model is trained to connect visual scenes and language instructions to manipulation actions, rather than merely describing an image.
  • Open research access: The checkpoint, source repository, and research documentation are publicly available, making it suitable for reproduction and experimentation.
  • Inspectable intermediate output: Depth and trajectory-related traces can help researchers examine what the model generated before an action is executed.
  • Adaptation potential: Its positioning as a pre-training checkpoint makes it useful for downstream mid-training on new robots, datasets, or task distributions.
  • Established multimodal foundation: Qwen2.5-7B and SigLIP2 provide a documented language-and-vision architecture rather than a robotics controller built only around low-level sensor inputs.

These are practical strengths for research teams that need an open starting point. They do not imply that the model is more accurate than every closed or task-specific robotics system, because the supplied research does not provide a direct comparative benchmark table.

Limitations and safety considerations

MolmoAct 7B-D Pretrain is specialized. It should not be treated as a general conversational assistant, a coding model, or a drop-in controller for arbitrary hardware. Its learned action space is bounded by its training data and by the robot embodiments, camera configurations, normalization statistics, and action representations used by the surrounding software.

The model's visual reasoning trace can be inspected, but inspection does not remove the risk of incorrect perception or unsafe movement. Any physical deployment should begin in simulation or under tightly controlled conditions, with emergency stops, motion limits, collision monitoring, human supervision, and safeguards appropriate to the robot. The model's output must be validated by a control system before it is allowed to move hardware.

The supplied research identifies the checkpoint as a preview release and states that it is licensed under Apache 2.0 for research and educational use. Researchers should still verify the current model-card terms, dataset conditions, and any project-specific restrictions before commercial or safety-critical deployment.

Pricing and access

There is no official hosted API price supplied for MolmoAct 7B-D Pretrain. The checkpoint is intended for self-hosted use through its open-weight distribution, so the direct model price is not a recurring per-token subscription. The actual cost depends on the researcher's hardware, storage, electricity, engineering time, and any external compute or hosting provider used to run inference.

This access model creates a different cost trade-off from a hosted robotics or multimodal API. Self-hosting can provide control over weights, inference code, and data handling, but it requires environment setup and compatible compute. The model record gives an editorial cost score of 9, but that score is a database evaluation rather than a provider-published price or benchmark and should not be interpreted as a guarantee of low operating cost.

Speed, coding, and tool-use trade-offs

MolmoAct is not primarily a software-coding model. Its coding score in the supplied record is an editorial rating of 2, not a provider claim, and the model's documented purpose is action reasoning for manipulation. Developers may still need ordinary programming skills to connect the checkpoint to cameras, simulation environments, robot controllers, safety systems, and action parsers.

Inference speed is also hardware-dependent. The record assigns a moderate editorial speed score of 4, but no authoritative latency benchmark is supplied. Casting FP32 weights to BF16 may improve memory and runtime characteristics on suitable hardware, although the resulting speed and action quality depend on the implementation and device. Because the model generates intermediate reasoning and trajectory information, a pipeline may need to balance deliberation time against the control loop's response requirements.

When to choose MolmoAct 7B-D Pretrain

Choose this checkpoint when you need an open research foundation for vision-language-action experiments and one or more of the following goals:

  • Reproducing the published zero-shot SimplerEnv experiments.
  • Continuing pre-training or mid-training on robotics data.
  • Studying how visual observations and language instructions can be converted into spatial action plans.
  • Adapting an open model to a new robot, camera arrangement, or manipulation dataset.
  • Inspecting intermediate depth and trajectory-related output before executing an action.

Another option may be more appropriate if you need a polished general assistant, built-in web access, ordinary function calling, guaranteed hosted latency, or a turnkey controller for a specific commercial robot. A task-specialized or fine-tuned MolmoAct release may also be a better starting point for a concrete deployment than this pre-training checkpoint. For safety-critical physical work, a validated conventional control stack with a narrowly tested perception model may be preferable to directly executing actions from a research checkpoint.

Bottom line

MolmoAct 7B-D Pretrain is most valuable as an open, adaptable research checkpoint for robotics action reasoning. Its combination of image input, language instructions, spatial traces, and parsed manipulation actions makes it relevant to researchers building or evaluating vision-language-action systems. Its limitations are equally important: there is no hosted API price or verified context limit, the 256-token example is not a confirmed maximum, tool calling is not established, and real-robot use requires substantial adaptation and safety engineering. For those who want to reproduce experiments or continue training an open robotics model, it is a targeted foundation; for general AI assistance or immediate hardware deployment, it is the wrong type of product.


Answers to Frequently Asked Questions

Where can I access the official MolmoAct 7B-D Pretrain checkpoint?
The canonical published checkpoint is available on Hugging Face as `allenai/MolmoAct-7B-D-Pretrain-0812`. It is intended for self-hosted use with compatible research code and environments, including reproducing zero-shot results in SimplerEnv.
Can MolmoAct 7B-D Pretrain control a physical robot out of the box?
No. It is a research and adaptation checkpoint, not a verified drop-in controller for arbitrary physical robots. Real-world deployment requires compatible action parsing, robot and camera adaptation, simulation or controlled testing, emergency stops, motion limits, collision monitoring, human supervision, and a safety-validated control system.
Which architecture does MolmoAct 7B-D Pretrain use?
MolmoAct 7B-D Pretrain is based on Qwen2.5-7B for language processing and uses SigLIP2 as its vision backbone. The published checkpoint is distributed in FP32 format, although the official inference example can cast the weights to BF16 on supported hardware.
What inputs and outputs does MolmoAct 7B-D Pretrain support?
The model accepts text instructions together with image observations. It generates a textual reasoning trace, depth-related information, trajectory information, and a robotics-oriented action representation that compatible MolmoAct code can parse into actions. Audio, video, generic tool calling, and distinct structured JSON output are not verified for this checkpoint.
What is MolmoAct 7B-D Pretrain designed for?
MolmoAct 7B-D Pretrain is an open-weight vision-language-action model from the Allen Institute for AI (Ai2) designed to connect visual observations and natural-language manipulation instructions to robot actions. It is intended primarily for robotics research, simulation, and downstream adaptation rather than as a general conversational assistant or turnkey robot controller.


Sources 3
Provider

About Allen Institute for Artificial Intelligence (Ai2)