MolmoMotion

MolmoMotion-FM

by Allen Institute for Artificial Intelligence (Ai2) · Documented research variant; no separately released public checkpoint identified in the current first-party model collection

Ai2’s MolmoMotion-FM is a specialized flow-matching model for forecasting continuous 3D trajectories of queried object points from RGB observations, language instructions, and initial 3D positions. It is designed to represent multiple plausible futures for robotics and motion-controlled video research. No separately downloadable FM checkpoint, hosted API pricing, standard context window, or general-purpose chat capability is established in the current first-party materials.

Reasoning Coding
MolmoMotion-FM is designed to answer a specialized question: given what a scene looks like, where selected points are located in 3D, and what action is intended, how might those points move next? The model uses a Molmo 2 vision-language backbone and a flow-matching trajectory generator to produce continuous 3D motion predictions. This makes it fundamentally different from a general-purpose chatbot or image generator. It is aimed at motion forecasting research, robotics, and explicit control of visual motion, with current availability and deployment details remaining limited.
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
4/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MolmoMotion
Model type Multimodal
Release date 2026-06-17
Status Documented research variant; no separately released public checkpoint identified in the current first-party model collection
Knowledge cutoff notes

No model-specific knowledge cutoff is published in the first-party MolmoMotion documentation. This is a motion-forecasting system rather than a conventional knowledge-grounded language model.

Model notes

MolmoMotion-FM is the flow-matching counterpart to MolmoMotion-AR. It predicts continuous 3D trajectories by transforming noise into motion and is intended to represent multiple plausible futures. The official project documentation describes the FM variant, while the current public MolmoMotion checkpoint collection and repository list released downloadable checkpoints as MolmoMotion-4B-H3-F30 and MolmoMotion-4B-H1-F32, both autoregressive. The exact FM parameter count, context window, maximum output-token limit, pricing, and standalone checkpoint identifier are not established by the current first-party release materials. Its output is a continuous 3D trajectory rather than ordinary text or an image, audio, or video file.

Model guide

MolmoMotion-FM: Flow-Matching 3D Motion Forecasting for Robotics and Video

MolmoMotion-FM is Ai2’s flow-matching variant of MolmoMotion, a research model that forecasts continuous future 3D trajectories for queried object points from short RGB observations, initial 3D positions, and natural-language action instructions. Its flow-matching design is intended to represent multiple plausible futures, making it relevant to robotics planning, language-guided motion prediction, and motion-controlled video generation. It is a documented research variant rather than a conventional hosted API model, and a separately downloadable FM checkpoint has not been identified in the current first-party release materials.

What is MolmoMotion-FM?

MolmoMotion-FM is the flow-matching variant of MolmoMotion, an open research system from the Allen Institute for AI (Ai2). It forecasts how selected points on objects may move through 3D space after receiving visual observations and a natural-language instruction describing an intended action.

For example, a system could provide a short sequence of RGB images, identify points on an object together with their initial 3D positions, and specify an instruction such as moving, opening, or manipulating the object. MolmoMotion-FM then predicts future 3D trajectories for those points in a metric camera-frame coordinate system. The result is motion data rather than an ordinary text response.

The model was introduced on June 17, 2026, alongside the autoregressive MolmoMotion-AR variant, the MolmoMotion-1M training corpus, and the PointMotionBench evaluation suite. These related releases position MolmoMotion-FM as a research component for language-guided 3D motion understanding and generation rather than as a general consumer assistant.

How the flow-matching design works

MolmoMotion-FM combines several kinds of information. Its inputs include a short history of RGB observations, a natural-language action description, 2D query-point features, and the initial 3D positions of the queried points. A Molmo 2 vision-language backbone connects visual content and language so that the requested action can be associated with the relevant objects and points.

The distinguishing part is the flow-matching trajectory generator. In accessible terms, it starts from noise and learns a continuous transformation that produces a future motion trajectory. This differs from an autoregressive model, which generates a sequence step by step, and from a system that represents coordinates as quantized text tokens.

The continuous formulation is intended to help the model express uncertainty. If an object could plausibly move in several different ways, a motion forecaster should not always be forced to select one apparently certain path. The supplied research describes MolmoMotion-FM as being designed to represent multiple plausible futures, although the available materials do not establish a standardized uncertainty metric or a published benchmark score for this specific variant.

Inputs and outputs

MolmoMotion-FM is multimodal in the sense that it consumes visual information together with language and 3D point information. Its documented inputs are:

  • Short RGB observation sequences.
  • A natural-language description of the intended action.
  • Selected 2D query points in the visual observations.
  • Initial 3D positions for the queried points.

Its output is a set of continuous future 3D trajectories for those points. These trajectories describe motion in 3D coordinate space; they are not generated images, videos, audio, or conversational text. The model can therefore serve as a motion-prediction component inside a larger robotics or visual-generation system, but it should not be evaluated as a standard text-generation model.

The supplied first-party research does not publish a conventional context-window size, maximum text-output limit, or token-based output specification. Those metrics are not directly applicable to the model’s primary trajectory output, and no verified values should be assumed.

What MolmoMotion-FM is used for

The model’s most suitable applications are those where predicted object motion is more useful than a textual explanation.

  • Language-guided 3D motion forecasting: Predict how points on rigid, articulated, or deformable objects may move after an instructed action.
  • Robotics planning: Supply a motion forecast that can help downstream systems reason about manipulation or possible object trajectories. The model is a forecasting component, not a complete robot controller.
  • Motion-conditioned video generation: Provide explicit 3D motion information to systems that generate or edit visual sequences according to a desired action.
  • Research into multimodal physical reasoning: Study how visual observations, language instructions, point tracking, and future-motion uncertainty can be combined.

These use cases depend on additional software and hardware around the model. MolmoMotion-FM does not, based on the supplied documentation, provide a complete hosted robotics stack, a web-search agent, or a turnkey video-generation product.

Where it fits in the MolmoMotion lineup

MolmoMotion-FM is the flow-matching counterpart to MolmoMotion-AR. The distinction is important: FM is described as transforming noise into continuous trajectories, while AR generates trajectory information autoregressively. The FM approach is especially relevant when representing several possible futures is a priority.

However, the current public release materials do not identify a separately downloadable MolmoMotion-FM checkpoint. The official repository and publicly listed MolmoMotion collection currently identify downloadable checkpoints as the autoregressive MolmoMotion-4B-H3-F30 and MolmoMotion-4B-H1-F32 models. As a result, MolmoMotion-FM should presently be treated as a documented research variant whose implementation or availability may not match the more directly accessible AR checkpoints.

This distinction matters for practical planning. A researcher looking for a model that can be downloaded and run immediately may need to evaluate the released autoregressive checkpoints instead, subject to their own documentation and license conditions. That does not make the FM formulation irrelevant; it means that the documented architecture and the currently identifiable public artifact are not necessarily the same thing.

Strengths and trade-offs

The main technical strength of MolmoMotion-FM is its focus. Rather than treating motion as a vague description in text, it predicts continuous 3D point trajectories conditioned on visual evidence and language. This makes its output more suitable for downstream geometric reasoning than a general vision-language model’s prose answer.

Its flow-matching formulation is also a useful fit for problems with ambiguous futures. A cup being moved, a deformable object being manipulated, or an articulated object changing configuration can involve more than one plausible trajectory. The model is designed with this uncertainty in mind instead of treating every prediction as a single deterministic text sequence.

The trade-off is specialization. The model is not documented as a general chat model, coding assistant, web-search system, speech model, image generator, or hosted API product. It also requires structured inputs such as query points and initial 3D positions, so it is not a drop-in replacement for a model that accepts only an image and a question.

Speed and cost should also be interpreted in deployment terms rather than token pricing. No provider token price, hosted inference price, or standard commercial service tier is identified. The supplied research gives no verified basis for comparing per-request cost with commercial APIs. Any actual cost will depend on the available checkpoint, hardware, implementation, and surrounding robotics or video pipeline.

Capabilities and unavailable specifications

AreaCurrent evidence
ProviderAllen Institute for AI (Ai2)
Model typeFlow-matching multimodal 3D motion forecaster
Visual inputRGB observations; video-like observation sequences are supported by the documented task
Language inputNatural-language action instructions
Primary outputContinuous future 3D point trajectories
General text outputNot the documented purpose
Tool or function callingNot identified
Web searchNot supported according to the supplied model record
PricingNo verified token or hosted-service pricing identified
Context lengthNot published for this variant
Maximum output tokensNot applicable or not published for the trajectory output
Standalone FM checkpointNot identified in the current first-party release materials

The absence of a published value should not be read as evidence that a capability exists or does not exist. In particular, the exact parameter count, supported sequence limits, maximum number of queried points, memory requirements, and inference speed for MolmoMotion-FM are not established by the supplied first-party sources.

Reasoning and coding capabilities

MolmoMotion-FM performs task-specific multimodal reasoning: it connects an action instruction with visual points and predicts their future motion. That is useful physical and geometric reasoning, but it should not be confused with open-ended chain-of-thought reasoning or broad knowledge-based problem solving.

Coding is not a model capability or output mode documented for MolmoMotion-FM. Developers may write code to use or adapt a research implementation, but the model itself is not presented as a code-generation assistant. Likewise, no tool-use or function-calling interface is identified.

When to choose MolmoMotion-FM

Choose MolmoMotion-FM when the central problem is forecasting object-point motion in 3D and the inputs naturally include visual observations, point locations, and a language instruction. It is particularly relevant for research comparing deterministic and multimodal motion forecasts, robotics systems that need a language-conditioned prediction stage, and video-generation workflows that benefit from explicit motion control.

Another motion model or the released MolmoMotion autoregressive variants may be more appropriate when immediate access to a public checkpoint is the priority. A general-purpose vision-language model is a better fit for image questions, document analysis, coding, or conversational assistance. A commercial video-generation platform is more appropriate when the goal is to produce a finished video rather than supply an intermediate 3D trajectory. A dedicated robotics planner or controller is still required when predictions must be converted into safe, executable actions.

Availability and bottom line

Ai2 publicly documents MolmoMotion-FM as part of the MolmoMotion research family, but the current first-party checkpoint listings identify released downloadable models as autoregressive variants rather than a separately downloadable FM checkpoint. There is no verified hosted API, subscription plan, token pricing, standard context window, or general-purpose chat interface for MolmoMotion-FM in the supplied materials.

Its value is therefore primarily research-oriented. MolmoMotion-FM offers a clearly defined approach to language-guided continuous 3D motion forecasting and is designed to model more than one plausible future. Its limitations are equally important: access to the specific FM artifact is unclear, deployment specifications are incomplete, and its output is specialized trajectory data rather than a ready-made answer, image, video, or robot action.


Answers to Frequently Asked Questions

Is a downloadable MolmoMotion-FM checkpoint available?
The current first-party release materials do not identify a separately downloadable MolmoMotion-FM checkpoint. They list the autoregressive MolmoMotion-4B-H3-F30 and MolmoMotion-4B-H1-F32 models as downloadable checkpoints, so researchers seeking immediate access may need to evaluate those variants instead.
What are the main use cases for MolmoMotion-FM?
Its primary use cases include language-guided 3D motion forecasting, robotics planning, motion-conditioned video generation, and research into multimodal physical reasoning. It provides motion forecasts for downstream systems rather than complete robot control, finished videos, or conversational answers.
How does MolmoMotion-FM differ from MolmoMotion-AR?
MolmoMotion-FM generates continuous trajectories by learning a transformation from noise, while MolmoMotion-AR produces trajectory information autoregressively, step by step. The flow-matching design is intended to represent multiple plausible future motions instead of selecting only one deterministic sequence.
What is MolmoMotion-FM designed to do?
MolmoMotion-FM is a flow-matching multimodal model from the Allen Institute for AI (Ai2) that forecasts future 3D trajectories for selected points on objects. It uses RGB observations, natural-language action instructions, 2D query points, and initial 3D positions to predict how those points may move.


Sources 5
Provider

About Allen Institute for Artificial Intelligence (Ai2)