MolmoPoint

MolmoPoint-Vid-4B

by Allen Institute for Artificial Intelligence (Ai2) · Current open-weight research model

An open-weight Ai2 vision-language model specialized for language-guided video pointing, object counting, temporal tracking, and spatiotemporal grounding. It returns decoded points, timestamps, and object identifiers rather than generated media, and is intended for local research inference rather than a managed commercial API.

Text Reasoning Coding
MolmoPoint-Vid-4B is a specialized video vision-language model from the Allen Institute for AI (Ai2). It accepts text instructions and video, then identifies where matching objects appear over time. Its grounding-token architecture is designed for precise pointing, counting, and tracking, making it more suitable for video understanding research and visual robotics than for ordinary chat or broad multimodal assistance.
Outputs

What MolmoPoint-Vid-4B can produce

Text
Inputs

What it can understand

Text Video Multimodal input
Model profile

Performance characteristics

4/10 Reasoning
2/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MolmoPoint
Model type Multimodal
Context window 35K tokens
Release date 2026-03-18
Status Current open-weight research model
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was published in the reviewed documentation. This is a video-grounding model whose primary behavior depends on supplied video and text inputs rather than a documented static knowledge date.

Model notes

Fully open Ai2 model released on March 18, 2026. The Hugging Face checkpoint is specialized for video pointing and is not intended as a generalist model. It generates grounding tokens that are decoded into timestamps, object IDs, and pixel coordinates rather than ordinary textual coordinates. The model card reports approximately 5B parameters and F32 weights. Video preprocessing supports up to 384 uniformly sampled frames at up to 2 FPS in the published configuration. The checkpoint is licensed under Apache 2.0, but training data may include datasets restricted to academic or non-commercial research use. The Hugging Face model card states that the checkpoint does not support training directly.

Model guide

MolmoPoint-Vid-4B: Ai2’s Open Model for Pointing at Objects in Video

MolmoPoint-Vid-4B is an open-weight Ai2 vision-language model built for language-guided video pointing. It grounds instructions in video by returning timestamps, object identifiers, and pixel coordinates for relevant objects, supporting counting and temporal tracking rather than general-purpose conversation.

What MolmoPoint-Vid-4B does

MolmoPoint-Vid-4B is an open-weight model designed to connect a natural-language instruction with specific locations in a video. A user might ask the model to identify every person wearing a particular color, count objects as they appear, or follow an object through several frames. The output is not simply a written description: it identifies points associated with timestamps and object identities.

This makes the model a grounding system. In this context, grounding means linking words such as “the red vehicle” or “each person entering the room” to concrete visual locations. The model is therefore aimed at tasks where knowing where and when something appears matters more than producing a fluent general answer.

Ai2 positions MolmoPoint-Vid-4B within its open Molmo research family. It is a specialized research checkpoint, not a unified consumer assistant or a hosted commercial API. The model is distributed through the allenai/MolmoPoint-Vid-4B Hugging Face repository.

How the grounding-token design works

Instead of generating ordinary textual coordinates directly, MolmoPoint-Vid-4B produces specialized grounding tokens. These tokens first select a coarse visual patch, refine that selection to a smaller subpatch, and then identify a location within the refined region. Decoding utilities convert the result into usable point data, including timestamps, object IDs, and pixel coordinates.

This approach is intended to make point prediction more efficient and more closely connected to the model’s visual representation. A special no-more-points class lets the model stop when it has identified all relevant objects rather than continuing to generate locations indefinitely.

For a developer, the distinction is important: loading the checkpoint and reading its generated text is not enough to obtain the final pointing results. Video preprocessing and the model’s point-extraction utilities are needed to translate the special tokens into coordinates and temporal metadata.

Main capabilities and practical use cases

  • Language-guided video pointing: locate objects described by a text instruction.
  • Object counting: identify multiple matching objects and return points for them.
  • Temporal tracking: associate points with object identifiers across video frames.
  • Spatiotemporal grounding: connect an instruction to both a position and a moment in the video.
  • Video understanding research: support experiments in tracking, visual reasoning, robotics, and grounded interaction.

A practical example would be asking the model to point to every cyclist visible in a clip and maintain their identities as they move. Another would be locating each package placed on a conveyor belt, with timestamps showing when the packages appear. These tasks are closer to visual measurement and annotation than to open-ended video conversation.

According to Ai2’s reported evaluation, MolmoPoint-Vid-4B achieved a 58.7 close-accuracy score on the Molmo2-VideoCount benchmark. This is a provider-reported benchmark result for the model’s intended video-counting task; it should not be interpreted as a general measure of conversational or multimodal performance.

Inputs, outputs, and processing limits

The documented input combination is text and video. Text provides the instruction, while video supplies the visual and temporal evidence. The model does not provide native image, audio, or video generation, and the supplied specifications do not identify audio input support.

The configuration lists a context length of 35,200 tokens. The available research does not specify a separate maximum output-token limit. Its effective video capacity is also affected by preprocessing: the published configuration supports up to 384 uniformly sampled frames at up to 2 frames per second. These settings determine how much of a source video is represented for inference and may affect temporal detail.

The primary result is text-compatible model output containing specialized grounding information that is decoded into points, timestamps, and object identifiers. It does not directly generate an image, edited video, audio file, or other media artifact. The model’s value lies in structured visual localization rather than media synthesis.

SpecificationVerified information
ProviderAllen Institute for AI (Ai2)
Model familyMolmoPoint
Model typeOpen-weight video vision-language and grounding model
Text inputYes
Video inputYes
Audio inputNot supported in the supplied specifications
Context length35,200 tokens
Published video preprocessingUp to 384 uniformly sampled frames at up to 2 FPS
Native media generationNo image, audio, or video output
Maximum output tokensNot specified

Implementation and availability

MolmoPoint-Vid-4B is intended for local or self-managed inference rather than a managed endpoint with provider billing. The public checkpoint is available under the Apache 2.0 license. However, the model documentation notes that some training data may be subject to academic, noncommercial, or other restrictions. The checkpoint’s license does not automatically remove obligations associated with third-party datasets or deployment contexts.

Local use requires a compatible environment, the model checkpoint, Ai2’s custom model code, and a video processor. The model card indicates that the Hugging Face checkpoint is intended for inference and does not support training directly through that checkpoint. Users who need to adapt the model should examine the project’s source code and documentation rather than assume that ordinary fine-tuning workflows will work unchanged.

Because this is an open research model, availability and implementation details can depend on the repository version, hardware, preprocessing configuration, and software dependencies. It should not be evaluated as if it offered the reliability, uptime guarantees, support contract, or standardized API behavior of a commercial hosted model.

Reasoning, coding, and tool support

MolmoPoint-Vid-4B performs a focused form of visual reasoning: it interprets a language instruction, identifies relevant visual content, and associates that content with locations across time. That does not make it a general reasoning model. The supplied specifications do not describe broad planning, long-form analysis, or independent world-knowledge reasoning as primary capabilities.

Coding is not a design goal. A program can use the model’s decoded points in a larger computer-vision or robotics pipeline, but the model itself is not intended for code generation or software development. Similarly, it has no documented web search, function-calling, plugin, or external tool-use capability. Any actions taken after a point is detected must be implemented by surrounding application code.

Speed, cost, and operational trade-offs

No input or output token prices are listed because MolmoPoint-Vid-4B is distributed as an open checkpoint rather than as a priced hosted inference service. Its financial cost depends on the hardware, storage, electricity, and infrastructure used to run it. Self-hosting can be attractive for research teams that need control over video data or repeated inference, but it transfers setup, optimization, scaling, and maintenance responsibilities to the user.

The model is approximately 5 billion parameters according to the model notes, with F32 weights mentioned in the published information. Video processing can also be computationally demanding, particularly when using many frames. In practice, a narrowly focused local model may be preferable to a larger general-purpose video model when point localization is the central task, but throughput will depend on the deployment hardware and preprocessing choices. The supplied research does not establish a universal latency comparison.

Strengths and limitations

Its main strength is specialization. MolmoPoint-Vid-4B produces a form of output that is directly useful for spatial and temporal annotation: points, timestamps, and object identities. The grounding-token mechanism also provides a purpose-built alternative to asking a general vision-language model to describe coordinates in free-form text. Its open weights and Apache 2.0 model license can support inspection, experimentation, and self-managed deployment.

The same specialization creates clear boundaries. It is not a general-purpose conversational vision model, and it is not designed for image or video generation. It does not include managed web access, function calling, audio processing, or a commercial API pricing layer. Users must handle model installation, video preparation, decoding, and infrastructure themselves. Results may also be affected by frame sampling, object ambiguity, camera motion, and the quality of the supplied instruction.

When to choose MolmoPoint-Vid-4B

Choose MolmoPoint-Vid-4B when the central requirement is to connect language with precise locations in video. It is a good candidate for research prototypes involving object counting, temporal tracking, grounded video analysis, robotics experiments, or automated annotation where open model artifacts and local control are important.

A different type of model may be more appropriate when the goal is general video question answering, extended conversational analysis, audio-visual understanding, code generation, media creation, or a production service with guaranteed availability and a simple API. A hosted commercial model may also be preferable when minimizing operational work is more important than controlling the checkpoint and processing data locally.

In short, MolmoPoint-Vid-4B should be evaluated as a specialized video pointing instrument. Its distinguishing output is not a polished chat response or generated media, but a set of grounded visual references that an application can use for measurement, tracking, and downstream action.


Answers to Frequently Asked Questions

What are the main limitations of MolmoPoint-Vid-4B?
MolmoPoint-Vid-4B is specialized for video pointing rather than general conversation, code generation, media creation, audio processing, or web-enabled tool use. Its effective video capacity depends on preprocessing, with the published configuration supporting up to 384 uniformly sampled frames at up to 2 frames per second. Users must also manage installation, decoding, hardware, and infrastructure themselves.
Where can developers access and run MolmoPoint-Vid-4B?
The checkpoint is available from the Hugging Face repository allenai/MolmoPoint-Vid-4B and is intended for local or self-managed inference. Running it requires a compatible environment, the checkpoint, Ai2’s custom model code, and a video processor.
What inputs and outputs does MolmoPoint-Vid-4B support?
The documented inputs are text instructions and video. Its output is structured grounding information that can be decoded into points, timestamps, and object identities. It does not natively generate images, audio, or video, and the supplied specifications do not identify audio input support.
What is MolmoPoint-Vid-4B designed to do?
MolmoPoint-Vid-4B is an open-weight video vision-language and grounding model from Ai2 that connects natural-language instructions to specific objects and locations across video frames. It can return points, timestamps, and object identifiers for tasks such as object counting, tracking, and spatiotemporal grounding.
How does MolmoPoint-Vid-4B identify objects in video?
The model generates specialized grounding tokens that select a coarse visual patch, refine it to a smaller region, and identify a location within that region. Decoding utilities then convert these tokens into pixel coordinates, timestamps, and object IDs.


Sources 6
Provider

About Allen Institute for Artificial Intelligence (Ai2)