Molmo-7B-D-0924
Local image understanding, visual question answering, captioning, document and chart analysis, counting, pointing, and multimodal research
Available as an open-weight downloadable checkpoint
Browse the AI models associated with Allen Institute for Artificial Intelligence (Ai2). Compare current and historical models by family, capabilities, context window, availability and intended use.
Local image understanding, visual question answering, captioning, document and chart analysis, counting, pointing, and multimodal research
Available as an open-weight downloadable checkpoint
Local image-and-text understanding, visual question answering, document and chart analysis, image captioning, and multimodal research
Available open-weight checkpoint; preview release
Self-hosted image understanding, visual question answering, captioning, image grounding, pointing, counting, and multimodal research.
Available open-weight legacy research model; newer Molmo 2 models are Ai2's current successor family.
Local image understanding, visual question answering, image captioning, document and chart analysis, counting, and experimentation with open multimodal models
Available open-weight preview checkpoint
Efficient local image and video understanding, visual grounding, pointing, captioning, counting, tracking, and multimodal research
Current open-weight model
Open multimodal research, video question answering, visual grounding, pointing, counting, captioning, object tracking, document understanding, and robotics perception
Current open-weight model; downloadable checkpoint
Open multimodal research, local image and video understanding, visual grounding, pointing, counting, captioning, and tracking
Current open-weight model
Open research and downstream fine-tuning for vision-guided robotic manipulation, spatial reasoning, trajectory planning, and robot action prediction
Available open-weight research model; superseded by MolmoAct2 as Ai2's newer MolmoAct generation
Robotic manipulation research, action reasoning, downstream mid-training, and reproducing zero-shot SimplerEnv experiments
Available open-weight preview checkpoint
Open robotics research, visual action reasoning, robot-manipulation experiments, and fine-tuning on custom robot datasets
Available open-weight preview checkpoint; intended for fine-tuning and downstream post-training
Open robotics research, embodied visual reasoning, robot-policy fine-tuning, and manipulation tasks on supported or closely related hardware
Current open-weight foundation checkpoint
Depth-aware robot manipulation research, embodied reasoning, and fine-tuning vision-language-action policies for target robot embodiments.
Current open-weight foundation checkpoint
Language-guided 3D trajectory forecasting, robotics planning research, and motion-conditioned video generation
Documented research variant; no separately released public checkpoint identified in the current first-party model collection
Open research and applications requiring image or video grounding, visual pointing, object localization, counting, tracking, spatial reasoning, and multimodal analysis.
Current open-weight model
GUI screenshot grounding, screen-element localization, computer-use perception, and research on pointing models
Current open-weight research model
Video object pointing, temporal grounding, object counting, tracking, and research on language-guided video understanding
Current open-weight research model
Open research, reproducible language-model experiments, local deployment, English text generation, and fine-tuning
Available open-weight model; not formally deprecated or retired
Local inference, open-model research, benchmarking, continued pretraining, and task-specific fine-tuning.
Available as a downloadable open-weight model; no official first-party hosted API identified.
Fully open language-model research, local inference, reproducible training experiments, evaluation, and downstream fine-tuning
Current downloadable open-weight model; no first-party hosted API deployment verified
Continued pretraining, supervised fine-tuning, reinforcement-learning research, language-model research, custom deployment, and applications requiring a transparent open training pipeline.
Current and downloadable open-weight model
Open, self-hosted chat assistants, instruction following, local coding help, tool-calling agents, and model research
Current open-weight model; available for download and self-hosted inference
Open research, mathematical reasoning, code generation, logic, long-context experimentation, and local self-hosted inference
Active open-weight model
Open-weight research, continued pretraining, fine-tuning, programming, mathematics, reading comprehension, and long-context language-model experiments
Current; open-weight and downloadable
Open-weight chat, tool-using assistants, multi-turn dialogue, synthetic data generation, self-hosting, and model research
Available open-weight checkpoint; superseded by Olmo 3.1 32B Instruct for newer instruction-tuned deployments
Open-model reasoning research, mathematics, coding, long-context analysis, and self-hosted deployment
Available as an open-weight model; superseded by the newer Olmo 3.1 32B Think model
Math-focused RLVR research, reinforcement-learning experiments, verifiable reward design, open-model evaluation, and further fine-tuning
Current open-weight research checkpoint
Mathematical reasoning, coding, complex multi-step tasks, open-model research, local deployment, and reinforcement-learning research
Current open-weight model
Open coding-model research, reinforcement learning from verifiable rewards, code-generation experiments, automated evaluation, and fine-tuning
Current, open-weight, experimental RL-Zero coding checkpoint
Self-hosted chat assistants, instruction following, tool-use agents, synthetic data, and open-model research
Current; open-weight and downloadable
English audio transcription, captions, meetings, lectures, podcasts, accessibility applications, and open ASR research
Current open-weight model
Open English speech-to-text transcription, captioning, meetings, lectures, calls, podcasts, and local ASR research
Available open-weight model
Open, self-hosted English transcription for meetings, lectures, calls, podcasts, accessibility, and speech-recognition research
Current open-weight model; downloadable and usable for local inference
English short- and long-form speech transcription, meeting and call transcription, lecture captioning, podcast processing, broadcast transcription, and speech analytics
Current open-weight model
Open English speech transcription, meeting and podcast transcription, captioning, timestamped audio indexing, and ASR research
Current open-weight model; available for download and self-managed inference
Efficient self-hosted English transcription, speech research, edge-oriented experiments, and applications where a small open ASR checkpoint is preferred.
Available open-weight checkpoint
Open-model research, efficient local text generation, language-model evaluation, and fine-tuning
Available open-weight release
Efficient satellite-image and Earth-observation embeddings, remote-sensing research, and downstream classification or segmentation
Current; downloadable open-weight research model
Satellite-image and Earth-observation embeddings, remote-sensing representation learning, geospatial classification, segmentation, and downstream fine-tuning
Current open-weight model
Efficient Earth observation embeddings, satellite image and time-series representation learning, geospatial classification, segmentation, and large-scale remote-sensing inference
Current open-weight model
Satellite-image embeddings, remote-sensing representation learning, geospatial segmentation, land-cover analysis, and Earth observation research
Current open-weight Earth observation foundation model
Open-weight language-model research, long-context text generation, local deployment, continued pretraining, and fine-tuning
Current open-weight model
Self-hosted instruction following, open post-training research, mathematics, coding, evaluation and domain adaptation
Available for download; superseded by Llama-3.1-Tulu-3.1-8B
Self-hosted instruction following, reasoning, mathematics, coding, open-model research, and reproducible post-training experiments
Available open-weight model
Large-scale research, instruction-following evaluation, open-weight post-training research, mathematical reasoning, coding benchmarks, and self-hosted experimentation.
Available as an open-weight research and educational model; no hosted inference provider deployment is currently listed on its Hugging Face model page.
Multimodal research, vision-language experiments, image generation, visual question answering, dense computer-vision tasks, and academic benchmarking
Open-weight research release; publicly available inference code and checkpoints; legacy relative to Ai2's newer multimodal models
Multimodal research, image understanding and generation, audio and video understanding, spatial prediction, embodied AI and robotic-manipulation experiments, and self-hosted academic prototyping.
Open-weight research release; publicly accessible checkpoints and source code; no official hosted inference API or commercial token pricing identified.
Open-vocabulary monocular 3D detection, spatial perception, robotics research, augmented reality, and lifting 2D prompts into metric 3D boxes
Current open-weight research model