MolmoPoint

MolmoPoint-8B

Open-weight Ai2 vision-language model for image, video, and multi-image understanding with specialized visual pointing and grounding capabilities.

Text Reasoning Coding
MolmoPoint-8B is an open-weight model from the Allen Institute for Artificial Intelligence (Ai2) that combines visual understanding with precise grounding. Instead of only describing what appears in an image or video, it can identify relevant locations using learned pointing tokens that are decoded into coordinates. The model is intended for local inference and research applications involving visual search, object localization, counting, spatial reasoning, and video tracking.
Outputs

What MolmoPoint-8B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Model profile

Performance characteristics

7/10 Reasoning
3/10 Coding
6/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MolmoPoint
Model type Multimodal
Release date 2026-03-18
Status Current open-weight model
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was identified for the exact MolmoPoint-8B checkpoint.

Model notes

MolmoPoint-8B is a fully open vision-language model developed by Ai2 and distributed through the allenai/MolmoPoint-8B Hugging Face checkpoint. It supports image, video, and multi-image understanding and grounding. Its pointing mechanism generates special grounding tokens that are decoded into image coordinates or video points using processor metadata and model extraction methods. The model has approximately 9B parameters and uses F32 safetensors in the published checkpoint. Ai2 reports 70.7% average accuracy on PointBench and 89.2 F1 on PixMo-Points. The Hugging Face checkpoint is intended for inference and does not support training directly; Ai2 provides separate training code. The model is licensed under Apache 2.0, but its training data includes third-party datasets with academic and non-commercial research restrictions. No official hosted API, token pricing, context window, maximum output limit, web-search integration, structured-output guarantee, or batch endpoint was identified.

Cost

Model pricing

Input No official hosted API pricing; downloadable checkpoint
Output No official hosted API pricing; downloadable checkpoint
Model guide

MolmoPoint-8B: Open Vision-Language Model for Visual Grounding

MolmoPoint-8B is an open-weight vision-language model from Ai2 built for image, video, and multi-image understanding with specialized visual pointing, object localization, counting, and tracking.

What is MolmoPoint-8B?

MolmoPoint-8B is a vision-language model developed by the Allen Institute for Artificial Intelligence, commonly known as Ai2. It belongs to Ai2's Molmo research line and is designed for tasks where a model must connect language with specific locations in visual content.

A conventional image-capable model might answer a question such as “Where is the red cup?” with a text description. MolmoPoint-8B is designed to go further: it can point to the relevant region. The same general approach can be used for identifying objects, counting instances, locating features, interpreting multiple images, and tracking visual targets through video.

The model was released on March 18, 2026, according to the supplied research. Its public Hugging Face checkpoint is distributed under the Apache 2.0 license and is intended for local inference rather than access through an official hosted API.

How its visual grounding mechanism works

MolmoPoint-8B's defining feature is its pointing mechanism. Rather than requiring the model to spell out coordinate values as ordinary text, the model generates special learned grounding tokens. Processing software then uses model metadata and extraction methods to decode those tokens into image coordinates or video points.

This design matters because coordinate generation is easy to make inconsistent when it is treated as ordinary language. A model may produce malformed values, use an unexpected format, or describe a location too generally. MolmoPoint-8B's specialized pointing tokens provide a model-specific representation for visual locations, making the output more suitable for applications that need to connect a response with a point or region in an image or video.

The model can support questions such as:

  • Which part of the image contains a specified object?
  • Where is an item located relative to other objects?
  • How many instances of a visual feature are present?
  • Which point in a video corresponds to a described object?
  • How does an object move across frames?

These examples describe the model's intended capabilities, not a guarantee that every image or video will be localized correctly. Accuracy depends on the visual content, prompt, preprocessing, and the surrounding inference implementation.

Supported inputs and outputs

The supplied model information identifies text, image, video, and multi-image understanding as supported inputs or use cases. Audio input is not identified as supported. MolmoPoint-8B produces text responses and grounding information; it is not an image, video, audio, or music generation model.

CapabilitySupported or documented status
Text inputSupported
Image inputSupported
Video inputSupported
Multiple-image understandingSupported
Audio inputNot identified
Text outputSupported
Visual pointing and groundingCore model capability
Image, video, or audio generationNot supported as a documented output type

The published information does not specify a context length, maximum output-token limit, streaming interface, tool-calling system, or guaranteed structured-output mode. Those omissions are important for production planning: developers should not assume that a hosted language-model API contract applies to this checkpoint.

Performance and model details

Ai2 reports 70.7% average accuracy on PointBench and 89.2 F1 on PixMo-Points for MolmoPoint-8B. These are provider-reported benchmark results, so they should be treated as measurements from Ai2's stated evaluation setup rather than universal guarantees. They are most useful for understanding the model's emphasis on pointing and grounding rather than as a complete measure of general visual intelligence.

The checkpoint contains approximately 9 billion parameters and uses F32 safetensors in the published repository. The model card is intended for inference; the checkpoint itself does not support training directly. Ai2 provides separate training code for users who need to investigate or adapt the model.

The model's published license is Apache 2.0. However, the supplied research notes that training data includes third-party datasets with academic and non-commercial research restrictions. Users should therefore distinguish between the license for the released model checkpoint and the terms governing particular training datasets or downstream uses.

Main strengths and trade-offs

Where MolmoPoint-8B is strong

  • Specialized visual localization: Pointing is a central design goal rather than an incidental feature.
  • Open local deployment: The downloadable checkpoint enables experimentation without depending on a proprietary hosted endpoint.
  • Multiple visual formats: The model is intended for images, videos, and multi-image understanding.
  • Useful spatial outputs: Decoded points can support interfaces or workflows that need a location, not just a prose explanation.
  • Research transparency: The public checkpoint, model documentation, benchmark reporting, and separate training code support reproducible investigation.

Important limitations

  • No official hosted API pricing or service: Users must plan for their own inference environment or use a third-party host if one is available.
  • Infrastructure burden: Running a roughly 9-billion-parameter F32 checkpoint locally can require substantial storage and compute resources. The supplied sources do not specify hardware requirements, so exact memory or latency should not be assumed.
  • Incomplete production interface: No official context window, maximum output limit, batch endpoint, web search, tool-use contract, or structured-output guarantee was identified.
  • Limited task scope: The model is not positioned as a general coding assistant, audio model, image generator, or broad productivity agent.
  • Training and data restrictions: The checkpoint's Apache 2.0 license does not remove restrictions that may apply to third-party training datasets.
  • Research-oriented availability: Access, documentation, and implementation details may change as Ai2 projects move between research and broader release stages.

Pricing and access

There is no official hosted API price listed for MolmoPoint-8B in the supplied research. The model is available as a downloadable checkpoint through the official allenai/MolmoPoint-8B Hugging Face repository, so the direct model cost is not expressed as a recurring per-token or per-request subscription price.

That does not mean inference is cost-free. Users may incur costs for GPUs, storage, electricity, cloud machines, orchestration, and engineering time. The financial advantage depends on usage volume and hardware availability. Occasional users may find a hosted multimodal service simpler, while teams with sustained workloads or strict control requirements may prefer operating an open checkpoint.

Reasoning, coding, and tool support

MolmoPoint-8B is primarily a multimodal perception and grounding model. Its useful reasoning capability is spatial and visual: interpreting a request, identifying relevant content, and associating that content with a point or location. The supplied editorial assessment rates its reasoning suitability at 7 out of 10, but that is an editorial score rather than an Ai2-published specification or benchmark.

Coding is not a core purpose. The supplied editorial assessment rates coding suitability at 3 out of 10, reflecting that the model is better suited to visual analysis than software development. It may be used inside a larger application that contains code, but there is no evidence here of a dedicated coding mode, programming benchmark, or general-purpose developer assistant interface.

Tool use and function calling are not documented for the checkpoint. Developers can build software around its outputs, including a pipeline that turns decoded points into interface markers or tracking instructions, but that surrounding orchestration should not be confused with native model tool support.

Speed and cost positioning

MolmoPoint-8B offers a different trade-off from a hosted, general-purpose multimodal assistant. Its open checkpoint provides control and avoids a mandatory provider API fee, but local inference transfers responsibility for hardware, optimization, scaling, and reliability to the operator.

The supplied editorial assessment gives the model a speed score of 6 out of 10 and a cost score of 9 out of 10. These scores are subjective evaluations, not measured service-level commitments. Actual performance will depend on the hardware, precision, video length, image resolution, batching, and implementation. Since the published information does not provide a context limit, maximum output size, or official latency figures, teams should benchmark their own workload before making deployment commitments.

When to choose MolmoPoint-8B

MolmoPoint-8B is a good candidate when visual location is more important than a polished hosted assistant experience. Consider it for:

  • Research on visual grounding and multimodal model behavior.
  • Applications that need to point to objects or regions in images.
  • Image counting and spatial question-answering experiments.
  • Video analysis involving object locations or movement across frames.
  • Multi-image comparisons where the model must relate language to visual evidence.
  • Private or customizable deployments where downloading and operating an open checkpoint is practical.

Another option may be more appropriate when the priority is a managed API, predictable latency, documented token limits, native function calling, web access, audio processing, general-purpose coding, or image generation. A smaller vision model may also be preferable when speed and hardware efficiency matter more than the specialized grounding mechanism. Conversely, a broader hosted multimodal model may be easier for a consumer-facing application that needs many capabilities in one service.

Bottom line

MolmoPoint-8B is best understood as a specialized open vision-language model for grounding visual answers in images and videos. Its distinguishing feature is the use of learned pointing tokens that can be decoded into visual coordinates, making it relevant to localization, counting, spatial reasoning, and tracking workflows. The model's open checkpoint and Apache 2.0 license support experimentation and local deployment, while the lack of an official hosted API, documented context and output limits, and built-in production tools creates additional engineering work. It is most compelling for researchers and developers who specifically need controllable visual pointing rather than a general-purpose multimodal assistant.


Answers to Frequently Asked Questions

What are the main limitations of MolmoPoint-8B?
MolmoPoint-8B requires local inference infrastructure, and the published information does not specify guaranteed hardware requirements, context limits, output limits, latency, streaming, structured output, or tool calling. It is specialized for visual perception and grounding rather than general coding, audio processing, image generation, or use as a broad hosted assistant. Third-party training data may also have restrictions separate from the model checkpoint's Apache 2.0 license.
Is MolmoPoint-8B available through an API, and what does it cost?
No official hosted API or recurring per-token price is listed in the supplied information. The model is available as a downloadable checkpoint from the official allenai/MolmoPoint-8B Hugging Face repository under the Apache 2.0 license. Users must account for their own GPU, storage, cloud, electricity, and engineering costs when running it.
What inputs and outputs does MolmoPoint-8B support?
The model supports text, images, video, and multiple-image understanding. It produces text responses and visual grounding information. Audio input, image generation, video generation, and audio generation are not documented as supported capabilities.
What is MolmoPoint-8B used for?
MolmoPoint-8B is an open vision-language model designed for visual grounding. It can identify objects or regions in images, answer spatial questions, count visual features, interpret multiple images, and track targets across video by connecting language with specific visual locations.
How does MolmoPoint-8B perform visual grounding?
Instead of generating coordinate values as ordinary text, MolmoPoint-8B produces learned grounding tokens. Inference software decodes these tokens using model metadata and extraction methods to obtain points or coordinates in images and videos.


Sources 2
Provider

About Allen Institute for Artificial Intelligence (Ai2)