What is MolmoMotion-FM?
MolmoMotion-FM is the flow-matching variant of MolmoMotion, an open research system from the Allen Institute for AI (Ai2). It forecasts how selected points on objects may move through 3D space after receiving visual observations and a natural-language instruction describing an intended action.
For example, a system could provide a short sequence of RGB images, identify points on an object together with their initial 3D positions, and specify an instruction such as moving, opening, or manipulating the object. MolmoMotion-FM then predicts future 3D trajectories for those points in a metric camera-frame coordinate system. The result is motion data rather than an ordinary text response.
The model was introduced on June 17, 2026, alongside the autoregressive MolmoMotion-AR variant, the MolmoMotion-1M training corpus, and the PointMotionBench evaluation suite. These related releases position MolmoMotion-FM as a research component for language-guided 3D motion understanding and generation rather than as a general consumer assistant.
How the flow-matching design works
MolmoMotion-FM combines several kinds of information. Its inputs include a short history of RGB observations, a natural-language action description, 2D query-point features, and the initial 3D positions of the queried points. A Molmo 2 vision-language backbone connects visual content and language so that the requested action can be associated with the relevant objects and points.
The distinguishing part is the flow-matching trajectory generator. In accessible terms, it starts from noise and learns a continuous transformation that produces a future motion trajectory. This differs from an autoregressive model, which generates a sequence step by step, and from a system that represents coordinates as quantized text tokens.
The continuous formulation is intended to help the model express uncertainty. If an object could plausibly move in several different ways, a motion forecaster should not always be forced to select one apparently certain path. The supplied research describes MolmoMotion-FM as being designed to represent multiple plausible futures, although the available materials do not establish a standardized uncertainty metric or a published benchmark score for this specific variant.
Inputs and outputs
MolmoMotion-FM is multimodal in the sense that it consumes visual information together with language and 3D point information. Its documented inputs are:
- Short RGB observation sequences.
- A natural-language description of the intended action.
- Selected 2D query points in the visual observations.
- Initial 3D positions for the queried points.
Its output is a set of continuous future 3D trajectories for those points. These trajectories describe motion in 3D coordinate space; they are not generated images, videos, audio, or conversational text. The model can therefore serve as a motion-prediction component inside a larger robotics or visual-generation system, but it should not be evaluated as a standard text-generation model.
The supplied first-party research does not publish a conventional context-window size, maximum text-output limit, or token-based output specification. Those metrics are not directly applicable to the model’s primary trajectory output, and no verified values should be assumed.
What MolmoMotion-FM is used for
The model’s most suitable applications are those where predicted object motion is more useful than a textual explanation.
- Language-guided 3D motion forecasting: Predict how points on rigid, articulated, or deformable objects may move after an instructed action.
- Robotics planning: Supply a motion forecast that can help downstream systems reason about manipulation or possible object trajectories. The model is a forecasting component, not a complete robot controller.
- Motion-conditioned video generation: Provide explicit 3D motion information to systems that generate or edit visual sequences according to a desired action.
- Research into multimodal physical reasoning: Study how visual observations, language instructions, point tracking, and future-motion uncertainty can be combined.
These use cases depend on additional software and hardware around the model. MolmoMotion-FM does not, based on the supplied documentation, provide a complete hosted robotics stack, a web-search agent, or a turnkey video-generation product.
Where it fits in the MolmoMotion lineup
MolmoMotion-FM is the flow-matching counterpart to MolmoMotion-AR. The distinction is important: FM is described as transforming noise into continuous trajectories, while AR generates trajectory information autoregressively. The FM approach is especially relevant when representing several possible futures is a priority.
However, the current public release materials do not identify a separately downloadable MolmoMotion-FM checkpoint. The official repository and publicly listed MolmoMotion collection currently identify downloadable checkpoints as the autoregressive MolmoMotion-4B-H3-F30 and MolmoMotion-4B-H1-F32 models. As a result, MolmoMotion-FM should presently be treated as a documented research variant whose implementation or availability may not match the more directly accessible AR checkpoints.
This distinction matters for practical planning. A researcher looking for a model that can be downloaded and run immediately may need to evaluate the released autoregressive checkpoints instead, subject to their own documentation and license conditions. That does not make the FM formulation irrelevant; it means that the documented architecture and the currently identifiable public artifact are not necessarily the same thing.
Strengths and trade-offs
The main technical strength of MolmoMotion-FM is its focus. Rather than treating motion as a vague description in text, it predicts continuous 3D point trajectories conditioned on visual evidence and language. This makes its output more suitable for downstream geometric reasoning than a general vision-language model’s prose answer.
Its flow-matching formulation is also a useful fit for problems with ambiguous futures. A cup being moved, a deformable object being manipulated, or an articulated object changing configuration can involve more than one plausible trajectory. The model is designed with this uncertainty in mind instead of treating every prediction as a single deterministic text sequence.
The trade-off is specialization. The model is not documented as a general chat model, coding assistant, web-search system, speech model, image generator, or hosted API product. It also requires structured inputs such as query points and initial 3D positions, so it is not a drop-in replacement for a model that accepts only an image and a question.
Speed and cost should also be interpreted in deployment terms rather than token pricing. No provider token price, hosted inference price, or standard commercial service tier is identified. The supplied research gives no verified basis for comparing per-request cost with commercial APIs. Any actual cost will depend on the available checkpoint, hardware, implementation, and surrounding robotics or video pipeline.
Capabilities and unavailable specifications
| Area | Current evidence |
|---|---|
| Provider | Allen Institute for AI (Ai2) |
| Model type | Flow-matching multimodal 3D motion forecaster |
| Visual input | RGB observations; video-like observation sequences are supported by the documented task |
| Language input | Natural-language action instructions |
| Primary output | Continuous future 3D point trajectories |
| General text output | Not the documented purpose |
| Tool or function calling | Not identified |
| Web search | Not supported according to the supplied model record |
| Pricing | No verified token or hosted-service pricing identified |
| Context length | Not published for this variant |
| Maximum output tokens | Not applicable or not published for the trajectory output |
| Standalone FM checkpoint | Not identified in the current first-party release materials |
The absence of a published value should not be read as evidence that a capability exists or does not exist. In particular, the exact parameter count, supported sequence limits, maximum number of queried points, memory requirements, and inference speed for MolmoMotion-FM are not established by the supplied first-party sources.
Reasoning and coding capabilities
MolmoMotion-FM performs task-specific multimodal reasoning: it connects an action instruction with visual points and predicts their future motion. That is useful physical and geometric reasoning, but it should not be confused with open-ended chain-of-thought reasoning or broad knowledge-based problem solving.
Coding is not a model capability or output mode documented for MolmoMotion-FM. Developers may write code to use or adapt a research implementation, but the model itself is not presented as a code-generation assistant. Likewise, no tool-use or function-calling interface is identified.
When to choose MolmoMotion-FM
Choose MolmoMotion-FM when the central problem is forecasting object-point motion in 3D and the inputs naturally include visual observations, point locations, and a language instruction. It is particularly relevant for research comparing deterministic and multimodal motion forecasts, robotics systems that need a language-conditioned prediction stage, and video-generation workflows that benefit from explicit motion control.
Another motion model or the released MolmoMotion autoregressive variants may be more appropriate when immediate access to a public checkpoint is the priority. A general-purpose vision-language model is a better fit for image questions, document analysis, coding, or conversational assistance. A commercial video-generation platform is more appropriate when the goal is to produce a finished video rather than supply an intermediate 3D trajectory. A dedicated robotics planner or controller is still required when predictions must be converted into safe, executable actions.
Availability and bottom line
Ai2 publicly documents MolmoMotion-FM as part of the MolmoMotion research family, but the current first-party checkpoint listings identify released downloadable models as autoregressive variants rather than a separately downloadable FM checkpoint. There is no verified hosted API, subscription plan, token pricing, standard context window, or general-purpose chat interface for MolmoMotion-FM in the supplied materials.
Its value is therefore primarily research-oriented. MolmoMotion-FM offers a clearly defined approach to language-guided continuous 3D motion forecasting and is designed to model more than one plausible future. Its limitations are equally important: access to the specific FM artifact is unclear, deployment specifications are incomplete, and its output is specialized trajectory data rather than a ready-made answer, image, video, or robot action.

