MolmoPoint

MolmoPoint-GUI-8B

by Allen Institute for Artificial Intelligence (Ai2) · Current open-weight research model

An open-weight GUI-specialized vision-language model from Ai2 that accepts a screenshot and instruction-like query, then returns a point identifying the requested interface element. It is designed for GUI grounding, screen understanding, and computer-use research rather than general autonomous action execution.

Text Reasoning Coding
MolmoPoint-GUI-8B is a specialized variant of MolmoPoint-8B trained for graphical-user-interface grounding. It accepts a single screenshot and an instruction-like text query, then identifies the requested interface element by producing a point through MolmoPoint grounding tokens.
Outputs

What MolmoPoint-GUI-8B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

4/10 Reasoning
2/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MolmoPoint
Model type Multimodal
Context window 37K tokens
Release date 2026-03-18
Status Current open-weight research model
Knowledge cutoff notes

Ai2's public model card and release materials do not state a knowledge cutoff for this exact model.

Model notes

MolmoPoint-GUI-8B is a fully open VLM developed by Ai2 and released under Apache 2.0. It is a GUI-specialized fine-tuning of MolmoPoint-8B and is trained for single-image GUI screenshot input with instruction-like queries. The model returns a single point identifying the requested interface element. It uses MolmoPoint grounding tokens rather than textual coordinate bins; generated tokens must be decoded with the supplied metadata and point-extraction utilities. The public Hugging Face checkpoint does not support training directly, while the Ai2 molmo2 repository provides GUI fine-tuning code. The 36,864-token sequence length is documented for the 8B MolmoPoint checkpoint conversion workflow and should be treated as the applicable model configuration. Ai2 reports 61.1 on ScreenSpot-Pro and 70.0 on OSWorldG, described as state-of-the-art among fully open models at release. No official hosted API pricing was identified because the model is distributed primarily as downloadable weights.

Model guide

MolmoPoint-GUI-8B: Open-Weight Model for Pointing to GUI Elements

MolmoPoint-GUI-8B is an open-weight Allen Institute for AI vision-language model specialized in locating and pointing to user-interface elements in single screenshots.

What is MolmoPoint-GUI-8B?

MolmoPoint-GUI-8B is an open-weight vision-language model from the Allen Institute for Artificial Intelligence (Ai2). Its purpose is narrower than that of a general conversational assistant: given a screenshot and a text instruction, it identifies the location of the requested interface element.

For example, a query might ask the model to point to a search box, a menu button, or a particular control visible on a screen. The model returns a single point identifying the target rather than carrying out the action itself. An external computer-use system could use that point as perception input before deciding whether and how to click, type, or perform another operation.

The model is a GUI-specialized fine-tuning of MolmoPoint-8B. It belongs to Ai2's open Molmo research family and is distributed primarily as downloadable model weights rather than as a conventional hosted consumer assistant or priced inference API.

How the model produces a GUI location

MolmoPoint-GUI-8B accepts one image, specifically a GUI screenshot, together with an instruction-like text query. It produces a point using MolmoPoint grounding tokens. These are model-specific output tokens designed to represent a location, rather than ordinary textual coordinate bins.

Using the result requires the decoding utilities and metadata supplied with the MolmoPoint implementation. In practical terms, developers should not assume that the raw generated token sequence is already a ready-to-use pair of pixel coordinates. The repository's point-extraction workflow is part of the intended integration path.

This design makes the model useful as a screen-understanding or perception component. It does not make the model a complete browser or desktop automation agent. A separate action layer is required to translate the predicted point into an interaction, and that layer would still need to manage application state, permissions, safety checks, and failures.

Verified specifications and supported modalities

SpecificationReported detail
ProviderAllen Institute for Artificial Intelligence (Ai2)
Model familyMolmoPoint
Model size8B-class checkpoint
Primary inputSingle image plus text query
Primary outputOne GUI grounding point and text-model output tokens used by the grounding workflow
LicenseApache 2.0
Documented sequence length36,864 tokens for the applicable 8B MolmoPoint conversion workflow
Official hosted priceNot identified; the model is primarily distributed as downloadable weights

The documented 36,864-token sequence length is associated with the 8B MolmoPoint checkpoint conversion workflow. It should be treated as the applicable model configuration, not as a promise of a hosted API context window. The supplied materials do not specify a separate maximum output-token limit for MolmoPoint-GUI-8B.

The supported input modalities are image and text. The research does not identify audio or video input for this model. It does not generate images, audio, or video: its relevant non-text output is a grounding point used to identify a screen location.

Main strengths

  • Focused GUI grounding: The model is trained for the specific task of locating interface elements, rather than being asked to perform GUI pointing as an incidental capability of a general assistant.
  • Open distribution: Ai2 provides the checkpoint under Apache 2.0, making it suitable for experimentation, inspection, and integration subject to the license and any applicable project requirements.
  • Useful perception primitive: A point prediction can serve as a compact input to a larger computer-use pipeline. This separates visual localization from the system that decides whether an action is appropriate.
  • Research and customization potential: The public molmo2 repository includes GUI fine-tuning code, while the model card and repository document the expected grounding-token and point-extraction workflow.
  • Reported benchmark position: Ai2 reports scores of 61.1 on ScreenSpot-Pro and 70.0 on OSWorldG, describing the results as state-of-the-art among fully open models at release. These are provider-reported claims, not independent evaluations established by the supplied research.

Its most important practical advantage is specialization. If the problem is “where is the control described by this instruction?” a dedicated pointing model may be easier to integrate and more predictable than using a general multimodal model for the same narrow task.

Limitations and trade-offs

MolmoPoint-GUI-8B does not provide end-to-end computer control. It identifies a point, but it does not independently browse, click, type, verify the result, recover from an incorrect action, or manage a desktop session. Those responsibilities belong to surrounding software.

The model is also not positioned as a general-purpose chat system. Its training and output format are aimed at screenshot grounding, so it is a poor fit for broad conversational assistance, long-form reasoning, general image generation, or workflows that require audio and video understanding.

Deployment requires more engineering than calling a typical hosted model endpoint. Users need to obtain and run the open checkpoint, follow the repository's conversion and inference instructions, and correctly decode the grounding output. The supplied research does not identify an official managed API, streaming interface, batch API, or provider-hosted pricing plan.

There are also operational trade-offs. Running an 8B open model locally or on privately managed infrastructure can provide control over deployment and data handling, but infrastructure, memory, latency, and maintenance costs depend on the chosen hardware and software stack. No universal speed or dollar-cost figure is supplied, so those factors should be measured in the target environment rather than inferred from the model name.

Reasoning, coding, and tool support

MolmoPoint-GUI-8B performs visual localization conditioned on an instruction. It should not be treated as a dedicated reasoning model with a separately documented chain-of-thought or advanced reasoning mode. Its useful “reasoning” behavior is task-specific: connecting the text request to the corresponding visible interface element.

It is not a coding model, although developers can use its output inside code for GUI automation or computer-use research. The model itself does not provide built-in function calling, tool execution, web search, browser control, or code execution according to the supplied specifications. Any such capabilities must be implemented by an external application.

Likewise, the point prediction is not the same as an action output. A point can tell another program where a target appears, but it does not authorize or perform the action. This distinction is important for safety-sensitive automation, where an external layer should confirm the target and apply appropriate constraints before interacting with a real application.

Pricing and access

No official hosted inference price was identified for MolmoPoint-GUI-8B. Ai2 distributes it primarily as downloadable open weights, so there is no verified per-input-token or per-output-token price to report. The practical cost is instead determined by storage, compute infrastructure, hosting, and engineering requirements.

The model is listed under the Apache 2.0 license in the supplied research. Developers should still review the current model card, repository instructions, and any associated dataset or project conditions before using it commercially or in a production system.

When to choose MolmoPoint-GUI-8B

Choose MolmoPoint-GUI-8B when the central requirement is locating a named or described interface element in a screenshot and you want an open model that can be inspected or integrated into your own perception pipeline. It is particularly relevant for:

  • GUI screenshot grounding and screen-element localization;
  • research on computer-use agents;
  • visual perception for browser or desktop automation prototypes;
  • experiments requiring open weights rather than a closed hosted service; and
  • systems where an external controller will handle actions after visual localization.

Another option may be more appropriate when you need a complete autonomous computer-use agent, a general conversational assistant, built-in tools, audio or video processing, image generation, or a managed API with predictable service-level and pricing characteristics. A general multimodal model may also be preferable when screenshot pointing is only one small part of a broader workflow and the surrounding tasks require extensive dialogue, planning, or document work.

Within the open-model ecosystem, MolmoPoint-GUI-8B is best understood as a specialized perception component rather than a complete automation product. Its value comes from doing one important task—mapping an instruction to a location on a GUI screenshot—in a form that another system can use.

Bottom line

MolmoPoint-GUI-8B is a focused, open-weight model for GUI grounding. It accepts a screenshot and instruction-like text, then returns a point identifying the requested interface element through MolmoPoint grounding tokens. The Apache 2.0 release and published fine-tuning and decoding resources make it attractive for research and custom computer-use pipelines.

Its limits are equally clear: it does not execute actions, provide a hosted API, or replace a general multimodal assistant. Teams choosing it should be prepared to manage model deployment, point decoding, action control, and evaluation themselves. For screenshot localization specifically, however, that specialization can be more useful than paying for or adapting a broader model whose main purpose lies elsewhere.


Answers to Frequently Asked Questions

When should developers choose MolmoPoint-GUI-8B?
Developers should choose it when they need specialized screenshot grounding, GUI element localization, or a visual perception component for a browser or desktop automation pipeline. A general multimodal model may be more suitable for broad conversation, built-in tools, audio or video processing, image generation, or complete autonomous computer control.
How is MolmoPoint-GUI-8B distributed and licensed?
MolmoPoint-GUI-8B is an open-weight model from the Allen Institute for Artificial Intelligence (Ai2), distributed primarily as downloadable weights under the Apache 2.0 license. No official hosted inference price or managed API was identified.
What inputs and outputs does MolmoPoint-GUI-8B support?
The model accepts a single GUI screenshot and an instruction-like text query. It produces a GUI grounding point using MolmoPoint grounding tokens, which must be decoded with the utilities and metadata provided by the MolmoPoint implementation.
What is MolmoPoint-GUI-8B used for?
MolmoPoint-GUI-8B is used to locate interface elements in GUI screenshots. Given a screenshot and a text instruction, it returns a point identifying the requested element, such as a search box, menu button, or other control.
Does MolmoPoint-GUI-8B perform clicks or control a computer?
No. MolmoPoint-GUI-8B only predicts the location of a target interface element. An external automation system must decode the point, verify it, and decide whether to click, type, or perform another action.


Sources 4
Provider

About Allen Institute for Artificial Intelligence (Ai2)