Molmo

Molmo-72B

by Allen Institute for Artificial Intelligence (Ai2) · Available open-weight legacy research model; newer Molmo 2 models are Ai2's current successor family.

Molmo-72B is Ai2’s largest original Molmo checkpoint. It combines Qwen2-72B with an OpenAI CLIP vision encoder to analyze images, answer visual questions, caption scenes, count objects, and provide grounding or pointing information. The model is available as Apache 2.0 open weights for local research and deployment, but its roughly 293 GB repository and 4,096-token context make it hardware-intensive. No official Ai2 hosted API price is documented.

Text Reasoning Coding
Molmo-72B is the largest checkpoint in Ai2’s original Molmo family. Released on September 24, 2024, it accepts text and images and produces text-based answers, descriptions, explanations, and visual grounding information. Its open weights, code, training materials, and PixMo data make it useful for researchers and organizations that need to inspect or run a multimodal model themselves. The trade-off is substantial: the roughly 293 GB model repository and 4,096-token configured context make it much more demanding to operate than smaller or hosted vision-language models.
Outputs

What Molmo-72B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
2/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Molmo
Model type Multimodal
Context window 4K tokens
Release date 2024-09-24
Status Available open-weight legacy research model; newer Molmo 2 models are Ai2's current successor family.
Knowledge cutoff notes

Ai2 does not publish a direct knowledge-cutoff date for the exact Molmo-72B-0924 checkpoint in the reviewed model documentation. The September 2024 suffix identifies the release checkpoint date, not a verified training-data cutoff.

Model notes

The canonical released checkpoint is allenai/Molmo-72B-0924. It is based on Qwen2-72B and uses an OpenAI CLIP ViT-L/14 336px vision encoder. The model accepts text and images and returns text, including visual grounding information. The Hugging Face repository is approximately 293 GB, making deployment hardware-intensive. The model is distributed as open weights under Apache 2.0 rather than through an official Ai2 token-priced API. The 4,096-token context value comes from the released model configuration. Fine-tuning is feasible because the weights are downloadable, but no official hosted fine-tuning service is documented. Scores are editorial comparative estimates, not vendor-provided ratings.

Model guide

Molmo-72B: Ai2’s Open-Weight Model for Deep Image Understanding

Molmo-72B is Ai2’s 72-billion-parameter open-weight vision-language model for image understanding, visual question answering, captioning, counting, grounding, and pointing. It combines a Qwen2-72B language model with an OpenAI CLIP vision encoder and is intended for self-hosted research and deployment rather than managed API use.

What is Molmo-72B?

Molmo-72B is a 72-billion-parameter vision-language model developed by the Allen Institute for AI, also known as Ai2. A vision-language model combines an image-processing component with a language model so it can interpret visual content and respond in text. Molmo-72B’s canonical released checkpoint is allenai/Molmo-72B-0924.

The model is designed primarily for image understanding rather than media generation. Given an image and a written prompt, it can answer questions, describe what it sees, interpret documents and charts, count objects, and identify or point to regions of an image. Its pointing and grounding behavior is represented through text-based coordinate information rather than a separately generated image or interactive visual overlay.

Molmo-72B was released on September 24, 2024, as part of Ai2’s first Molmo model family. That family also included smaller or differently configured checkpoints such as MolmoE-1B, Molmo-7B-O, and Molmo-7B-D. Ai2’s current Molmo materials emphasize the newer Molmo 2 family, so Molmo-72B is best understood as an available legacy research checkpoint rather than Ai2’s current flagship multimodal model.

Architecture and training approach

Molmo-72B combines a Qwen2-72B language backbone with an OpenAI CLIP ViT-L/14 336px vision encoder. In practical terms, the vision encoder converts image content into representations that the language model can use, while the language component produces the final textual response.

The released architecture uses multiscale, multi-crop image preprocessing. This means the input image can be examined at different scales and crops before visual information is passed into the language model. The system also includes a connector that maps visual features into the language model’s representation space and a decoder-only transformer language model for generating responses.

Ai2’s published training description identifies two main stages: multimodal pretraining focused on caption generation, followed by supervised fine-tuning. The supplied research states that the training recipe did not use RLHF. Training used Ai2’s PixMo datasets, including data for detailed image captioning, visual question answering, pointing, counting, documents, and clocks. These materials are important to Molmo-72B’s focus on detailed visual interpretation rather than text-only generation.

Inputs, outputs, and supported capabilities

Molmo-72B accepts text and images. It returns text, including ordinary natural-language answers and text representations of visual locations. It does not natively generate images, audio, or video.

  • Text input: Supported.
  • Image input: Supported.
  • Audio input: Not supported according to the supplied model specification.
  • Video input: Not supported according to the supplied model specification.
  • Text output: Supported.
  • Image, audio, video, and music output: Not supported.

Typical uses include asking what is shown in a photograph, extracting information from a chart, describing a document page, counting visible objects, or asking the model to identify where an object appears. Captioning and visual question answering are central use cases. Grounding and pointing are also notable because the model can communicate approximate visual locations instead of only describing an object in prose.

There is no verified provider-published maximum output-token value in the supplied research. The released configuration specifies a 4,096-token maximum position length. That context value applies to the model’s configured input-and-generation sequence capacity and should not be presented as a separate guaranteed output allowance.

Reasoning, coding, and tool support

Molmo-72B can perform useful visual reasoning in the practical sense of combining image evidence with a written question. For example, it can connect a question about a chart or document to visible information and return an explanation. However, the supplied research does not identify a separate chain-of-thought mode, formal reasoning mode, or provider-defined reasoning control. Its editorial reasoning score is therefore not a vendor specification and should not be treated as a benchmark result.

Coding is not the model’s primary purpose. Because it produces text, it can potentially describe code shown in an image or generate code-like text in response to a prompt, but the supplied research does not document a specialized coding capability or coding benchmark. It should not be selected over a dedicated coding model simply because its underlying language model is large.

The model does not have documented native web search, function calling, or agent tools. It is distributed as downloadable weights and inference code rather than as a unified assistant with built-in browsing or external actions. Developers could build surrounding software that sends its text output to other systems, but that would be application-level integration, not a verified native Molmo-72B tool feature.

Deployment and operational requirements

Molmo-72B is intended for users who can manage substantial model infrastructure. The Hugging Face repository is approximately 293 GB, making the checkpoint far too large for many ordinary laptops or single modest GPUs. Practical inference generally requires significant GPU memory, quantization, model sharding, or hosted infrastructure.

The model can be loaded locally with the Transformers library and Ai2’s custom model code. It can also be served through compatible inference systems such as vLLM, subject to the serving system’s support and the operator’s configuration. Because the weights are downloadable, organizations can inspect, adapt, quantize, and fine-tune the model within the conditions of its license and their own operational setup.

This deployment model gives Molmo-72B a different profile from a hosted commercial vision API. Self-hosting can provide more control over data handling, versioning, infrastructure, and customization, but the operator must supply the hardware, software maintenance, monitoring, scaling, and performance engineering. There is no official Ai2 token-priced managed API documented for this checkpoint.

Pricing, license, and availability

There is no verified per-token or subscription price for Molmo-72B in the supplied research. Ai2 distributes the checkpoint as downloadable open weights rather than through an official priced inference API. That does not mean deployment is cost-free: users may incur expenses for GPUs, storage, electricity, cloud instances, engineering, and maintenance.

The model weights and supporting artifacts are distributed under the Apache 2.0 license according to the supplied research. Users should still review the exact repository terms, associated dataset conditions, and any responsible-use requirements before commercial or high-stakes deployment.

Molmo-72B remains available under the checkpoint name allenai/Molmo-72B-0924. Ai2’s emphasis has shifted toward Molmo 2, which is the more relevant family to investigate when a newer Ai2 multimodal model is required. Molmo-72B remains useful when reproducibility, access to the original open checkpoint, or compatibility with an existing Molmo-based research workflow matters more than using the newest model.

Strengths and limitations

Where Molmo-72B is strong

  • Open distribution: Downloadable weights and supporting materials enable inspection and self-hosted experimentation.
  • Detailed image understanding: Its training focus covers captioning, visual question answering, documents, counting, pointing, and grounding.
  • Visual location output: Text-based coordinates can be useful in research involving object localization or image interaction.
  • Research transparency: Ai2 provides model, code, training, and dataset-related materials intended to support reproducible research.
  • Customization: Local operators can investigate quantization, adaptation, fine-tuning, and deployment configurations instead of relying only on a fixed hosted endpoint.

Where it is limited

  • Very high hardware demand: The approximately 293 GB repository makes local operation difficult for many teams.
  • Short configured context: The released configuration specifies 4,096 tokens, which limits long multimodal conversations and large document workflows.
  • No official managed API pricing: Developers must operate the model themselves or find compatible external infrastructure.
  • No native media generation: It does not generate images, audio, or video.
  • No documented native tool layer: Web search, function calling, and external actions are not identified as built-in capabilities.
  • Older family position: Later Molmo 2 models have superseded it in Ai2’s current multimodal lineup and reportedly improve on some visual grounding and video-related tasks.

When to choose Molmo-72B

Choose Molmo-72B when the main requirement is open, self-hosted image understanding and your team can support a very large checkpoint. It is a reasonable candidate for research on visual question answering, image captioning, document or chart interpretation, counting, grounding, and pointing. It is also attractive when access to model artifacts and the ability to inspect or modify the system are more important than turnkey hosting.

Compared with a small hosted vision model, Molmo-72B favors control and model scale over speed, simplicity, and infrastructure cost. A hosted service or smaller checkpoint may be more appropriate for interactive applications that need low latency, predictable per-request pricing, automatic scaling, or minimal operations work. A newer Molmo 2 checkpoint may be preferable when current Ai2 model development, improved visual grounding, or video-related capabilities are priorities, although the supplied research does not provide a detailed feature-by-feature specification for those models.

Molmo-72B is a poor fit for applications that need image generation, speech, video generation, built-in web browsing, long-context document analysis, or a supported production API with published service-level guarantees. It is also not the natural choice for a dedicated coding workflow. In those cases, a model specifically designed for the required modality, context length, coding task, or hosted operating model may be more suitable.

Bottom line

Molmo-72B is a large, open-weight vision-language model whose main value lies in transparent, self-hosted image understanding. Its combination of Qwen2-72B, a CLIP vision encoder, PixMo training data, and visual grounding support makes it relevant for multimodal research and specialized deployments. The same design also creates its central drawback: operating a roughly 293 GB checkpoint with a 4,096-token context requires serious infrastructure. For teams that want an inspectable Ai2 model and can absorb that cost, Molmo-72B remains a useful research checkpoint. For most low-latency, low-maintenance, or current-production workloads, a smaller, newer, or hosted alternative is likely to be more practical.


Answers to Frequently Asked Questions

When should I choose Molmo-72B instead of a newer or hosted vision model?
Choose Molmo-72B when open weights, self-hosting, inspectability, customization, and detailed image understanding are more important than low latency or simple operations. It is suitable for research involving visual question answering, captioning, documents, charts, counting, grounding, and pointing. A smaller hosted model or a newer Molmo 2 checkpoint may be better for automatic scaling, predictable pricing, long-context workflows, video capabilities, or current production deployments.
Is Molmo-72B free to use, and what license does it have?
Ai2 distributes Molmo-72B as downloadable open weights rather than through an official token-priced inference API. The model weights and supporting artifacts are distributed under the Apache 2.0 license according to the supplied research. Deployment is not cost-free because users must cover hardware, storage, cloud, electricity, engineering, and maintenance expenses, and should review the repository terms before use.
How much hardware is required to run Molmo-72B?
Molmo-72B requires substantial infrastructure. Its Hugging Face repository is approximately 293 GB, making it impractical for many ordinary laptops or single modest GPUs. Deployment generally requires significant GPU memory, quantization, model sharding, or hosted infrastructure. It can be loaded with Transformers and may be served through compatible systems such as vLLM.
What is Molmo-72B?
Molmo-72B is a 72-billion-parameter open-weight vision-language model developed by the Allen Institute for AI (Ai2). Its canonical checkpoint is allenai/Molmo-72B-0924, and it is designed for image understanding tasks such as visual question answering, image captioning, document and chart interpretation, counting, pointing, and grounding.
What can Molmo-72B do, and what inputs does it support?
Molmo-72B accepts text and images and returns text, including natural-language answers and text-based visual coordinates. It can describe images, answer questions about visual content, interpret documents and charts, count objects, and identify approximate object locations. It does not natively support audio or video input, or image, audio, video, or music generation.


Sources 6
Provider

About Allen Institute for Artificial Intelligence (Ai2)