What MolmoPoint-Vid-4B does
MolmoPoint-Vid-4B is an open-weight model designed to connect a natural-language instruction with specific locations in a video. A user might ask the model to identify every person wearing a particular color, count objects as they appear, or follow an object through several frames. The output is not simply a written description: it identifies points associated with timestamps and object identities.
This makes the model a grounding system. In this context, grounding means linking words such as “the red vehicle” or “each person entering the room” to concrete visual locations. The model is therefore aimed at tasks where knowing where and when something appears matters more than producing a fluent general answer.
Ai2 positions MolmoPoint-Vid-4B within its open Molmo research family. It is a specialized research checkpoint, not a unified consumer assistant or a hosted commercial API. The model is distributed through the allenai/MolmoPoint-Vid-4B Hugging Face repository.
How the grounding-token design works
Instead of generating ordinary textual coordinates directly, MolmoPoint-Vid-4B produces specialized grounding tokens. These tokens first select a coarse visual patch, refine that selection to a smaller subpatch, and then identify a location within the refined region. Decoding utilities convert the result into usable point data, including timestamps, object IDs, and pixel coordinates.
This approach is intended to make point prediction more efficient and more closely connected to the model’s visual representation. A special no-more-points class lets the model stop when it has identified all relevant objects rather than continuing to generate locations indefinitely.
For a developer, the distinction is important: loading the checkpoint and reading its generated text is not enough to obtain the final pointing results. Video preprocessing and the model’s point-extraction utilities are needed to translate the special tokens into coordinates and temporal metadata.
Main capabilities and practical use cases
- Language-guided video pointing: locate objects described by a text instruction.
- Object counting: identify multiple matching objects and return points for them.
- Temporal tracking: associate points with object identifiers across video frames.
- Spatiotemporal grounding: connect an instruction to both a position and a moment in the video.
- Video understanding research: support experiments in tracking, visual reasoning, robotics, and grounded interaction.
A practical example would be asking the model to point to every cyclist visible in a clip and maintain their identities as they move. Another would be locating each package placed on a conveyor belt, with timestamps showing when the packages appear. These tasks are closer to visual measurement and annotation than to open-ended video conversation.
According to Ai2’s reported evaluation, MolmoPoint-Vid-4B achieved a 58.7 close-accuracy score on the Molmo2-VideoCount benchmark. This is a provider-reported benchmark result for the model’s intended video-counting task; it should not be interpreted as a general measure of conversational or multimodal performance.
Inputs, outputs, and processing limits
The documented input combination is text and video. Text provides the instruction, while video supplies the visual and temporal evidence. The model does not provide native image, audio, or video generation, and the supplied specifications do not identify audio input support.
The configuration lists a context length of 35,200 tokens. The available research does not specify a separate maximum output-token limit. Its effective video capacity is also affected by preprocessing: the published configuration supports up to 384 uniformly sampled frames at up to 2 frames per second. These settings determine how much of a source video is represented for inference and may affect temporal detail.
The primary result is text-compatible model output containing specialized grounding information that is decoded into points, timestamps, and object identifiers. It does not directly generate an image, edited video, audio file, or other media artifact. The model’s value lies in structured visual localization rather than media synthesis.
| Specification | Verified information |
|---|---|
| Provider | Allen Institute for AI (Ai2) |
| Model family | MolmoPoint |
| Model type | Open-weight video vision-language and grounding model |
| Text input | Yes |
| Video input | Yes |
| Audio input | Not supported in the supplied specifications |
| Context length | 35,200 tokens |
| Published video preprocessing | Up to 384 uniformly sampled frames at up to 2 FPS |
| Native media generation | No image, audio, or video output |
| Maximum output tokens | Not specified |
Implementation and availability
MolmoPoint-Vid-4B is intended for local or self-managed inference rather than a managed endpoint with provider billing. The public checkpoint is available under the Apache 2.0 license. However, the model documentation notes that some training data may be subject to academic, noncommercial, or other restrictions. The checkpoint’s license does not automatically remove obligations associated with third-party datasets or deployment contexts.
Local use requires a compatible environment, the model checkpoint, Ai2’s custom model code, and a video processor. The model card indicates that the Hugging Face checkpoint is intended for inference and does not support training directly through that checkpoint. Users who need to adapt the model should examine the project’s source code and documentation rather than assume that ordinary fine-tuning workflows will work unchanged.
Because this is an open research model, availability and implementation details can depend on the repository version, hardware, preprocessing configuration, and software dependencies. It should not be evaluated as if it offered the reliability, uptime guarantees, support contract, or standardized API behavior of a commercial hosted model.
Reasoning, coding, and tool support
MolmoPoint-Vid-4B performs a focused form of visual reasoning: it interprets a language instruction, identifies relevant visual content, and associates that content with locations across time. That does not make it a general reasoning model. The supplied specifications do not describe broad planning, long-form analysis, or independent world-knowledge reasoning as primary capabilities.
Coding is not a design goal. A program can use the model’s decoded points in a larger computer-vision or robotics pipeline, but the model itself is not intended for code generation or software development. Similarly, it has no documented web search, function-calling, plugin, or external tool-use capability. Any actions taken after a point is detected must be implemented by surrounding application code.
Speed, cost, and operational trade-offs
No input or output token prices are listed because MolmoPoint-Vid-4B is distributed as an open checkpoint rather than as a priced hosted inference service. Its financial cost depends on the hardware, storage, electricity, and infrastructure used to run it. Self-hosting can be attractive for research teams that need control over video data or repeated inference, but it transfers setup, optimization, scaling, and maintenance responsibilities to the user.
The model is approximately 5 billion parameters according to the model notes, with F32 weights mentioned in the published information. Video processing can also be computationally demanding, particularly when using many frames. In practice, a narrowly focused local model may be preferable to a larger general-purpose video model when point localization is the central task, but throughput will depend on the deployment hardware and preprocessing choices. The supplied research does not establish a universal latency comparison.
Strengths and limitations
Its main strength is specialization. MolmoPoint-Vid-4B produces a form of output that is directly useful for spatial and temporal annotation: points, timestamps, and object identities. The grounding-token mechanism also provides a purpose-built alternative to asking a general vision-language model to describe coordinates in free-form text. Its open weights and Apache 2.0 model license can support inspection, experimentation, and self-managed deployment.
The same specialization creates clear boundaries. It is not a general-purpose conversational vision model, and it is not designed for image or video generation. It does not include managed web access, function calling, audio processing, or a commercial API pricing layer. Users must handle model installation, video preparation, decoding, and infrastructure themselves. Results may also be affected by frame sampling, object ambiguity, camera motion, and the quality of the supplied instruction.
When to choose MolmoPoint-Vid-4B
Choose MolmoPoint-Vid-4B when the central requirement is to connect language with precise locations in video. It is a good candidate for research prototypes involving object counting, temporal tracking, grounded video analysis, robotics experiments, or automated annotation where open model artifacts and local control are important.
A different type of model may be more appropriate when the goal is general video question answering, extended conversational analysis, audio-visual understanding, code generation, media creation, or a production service with guaranteed availability and a simple API. A hosted commercial model may also be preferable when minimizing operational work is more important than controlling the checkpoint and processing data locally.
In short, MolmoPoint-Vid-4B should be evaluated as a specialized video pointing instrument. Its distinguishing output is not a polished chat response or generated media, but a set of grounded visual references that an application can use for measurement, tracking, and downstream action.

