Model catalog

Allen Institute for Artificial Intelligence (Ai2) Models

Browse the AI models associated with Allen Institute for Artificial Intelligence (Ai2). Compare current and historical models by family, capabilities, context window, availability and intended use.

48 models tracked
48 Total models
17 Model families
7 Model types
46 Current / accessible
All models

Allen Institute for Artificial Intelligence (Ai2) model catalog

Local image understanding, visual question answering, captioning, document and chart analysis, counting, pointing, and multimodal research

Type Multimodal
Context 4K
Reasoning 6/10
Speed 6/10
Multimodal Image input Fine-tuning
Status

Available as an open-weight downloadable checkpoint

Input No official hosted API pricing; self-hosted weights
Output No official hosted API pricing; self-hosted weights
View model →

Local image-and-text understanding, visual question answering, document and chart analysis, image captioning, and multimodal research

Type Multimodal
Context 4K
Reasoning 6/10
Speed 6/10
Multimodal Image input Fine-tuning
Status

Available open-weight checkpoint; preview release

Input No official hosted API pricing; downloadable weights for self-hosted deployment
Output No official hosted API pricing; deployment cost depends on self-hosted or third-party infrastructure
View model →

Self-hosted image understanding, visual question answering, captioning, image grounding, pointing, counting, and multimodal research.

Type Multimodal
Context 4K
Reasoning 7/10
Speed 2/10
Multimodal Image input Fine-tuning
Status

Available open-weight legacy research model; newer Molmo 2 models are Ai2's current successor family.

View model →

Local image understanding, visual question answering, image captioning, document and chart analysis, counting, and experimentation with open multimodal models

Type Multimodal
Context 4K
Reasoning 5/10
Speed 8/10
Multimodal Image input
Status

Available open-weight preview checkpoint

View model →
Allen Institute for Artificial Intelligence (Ai2) logo
Molmo 2

Molmo2-4B

Efficient local image and video understanding, visual grounding, pointing, captioning, counting, tracking, and multimodal research

Type Multimodal
Reasoning 6/10
Speed 8/10
Multimodal Image input Video input
Status

Current open-weight model

Input No official hosted API pricing; self-hosted model weights
Output No official hosted API pricing; self-hosted model weights
View model →
Allen Institute for Artificial Intelligence (Ai2) logo
Molmo 2

Molmo2-8B

Open multimodal research, video question answering, visual grounding, pointing, counting, captioning, object tracking, document understanding, and robotics perception

Type Multimodal
Context 37K
Reasoning 7/10
Speed 6/10
Multimodal Image input Video input
Status

Current open-weight model; downloadable checkpoint

View model →

Open multimodal research, local image and video understanding, visual grounding, pointing, counting, captioning, and tracking

Type Multimodal
Context 66K
Reasoning 7/10
Speed 6/10
Multimodal Image input Video input
Status

Current open-weight model

Input Not applicable; open-weight model with no official hosted API token pricing found
Output Not applicable; open-weight model with no official hosted API token pricing found
View model →

Open research and downstream fine-tuning for vision-guided robotic manipulation, spatial reasoning, trajectory planning, and robot action prediction

Type Multimodal
Context 4K
Reasoning 7/10
Speed 5/10
Multimodal Image input Media output
Status

Available open-weight research model; superseded by MolmoAct2 as Ai2's newer MolmoAct generation

View model →

Robotic manipulation research, action reasoning, downstream mid-training, and reproducing zero-shot SimplerEnv experiments

Type Multimodal
Reasoning 7/10
Speed 4/10
Multimodal Image input Fine-tuning
Status

Available open-weight preview checkpoint

Input No official hosted API pricing; self-hosted checkpoint
Output No official hosted API pricing; self-hosted checkpoint
View model →

Open robotics research, visual action reasoning, robot-manipulation experiments, and fine-tuning on custom robot datasets

Type Multimodal
Reasoning 7/10
Speed 4/10
Multimodal Image input Media output
Status

Available open-weight preview checkpoint; intended for fine-tuning and downstream post-training

View model →
Allen Institute for Artificial Intelligence (Ai2) logo
MolmoAct2

MolmoAct2

Open robotics research, embodied visual reasoning, robot-policy fine-tuning, and manipulation tasks on supported or closely related hardware

Type Multimodal
Context 16K
Reasoning 8/10
Speed 8/10
Multimodal Image input Fine-tuning
Status

Current open-weight foundation checkpoint

View model →

Depth-aware robot manipulation research, embodied reasoning, and fine-tuning vision-language-action policies for target robot embodiments.

Type Robotics
Reasoning 7/10
Speed 5/10
Multimodal Image input Fine-tuning
Status

Current open-weight foundation checkpoint

View model →
Allen Institute for Artificial Intelligence (Ai2) logo
MolmoMotion

MolmoMotion-FM

Language-guided 3D trajectory forecasting, robotics planning research, and motion-conditioned video generation

Type Multimodal
Reasoning 2/10
Speed 4/10
Multimodal Image input Video input
Status

Documented research variant; no separately released public checkpoint identified in the current first-party model collection

View model →
Allen Institute for Artificial Intelligence (Ai2) logo
MolmoPoint

MolmoPoint-8B

Open research and applications requiring image or video grounding, visual pointing, object localization, counting, tracking, spatial reasoning, and multimodal analysis.

Type Multimodal
Reasoning 7/10
Speed 6/10
Multimodal Image input Video input
Status

Current open-weight model

Input No official hosted API pricing; downloadable checkpoint
Output No official hosted API pricing; downloadable checkpoint
View model →

GUI screenshot grounding, screen-element localization, computer-use perception, and research on pointing models

Type Multimodal
Context 37K
Reasoning 4/10
Speed 5/10
Multimodal Image input Fine-tuning
Status

Current open-weight research model

View model →

Video object pointing, temporal grounding, object counting, tracking, and research on language-guided video understanding

Type Multimodal
Context 35K
Reasoning 4/10
Speed 5/10
Multimodal Video input
Status

Current open-weight research model

View model →

Open research, reproducible language-model experiments, local deployment, English text generation, and fine-tuning

Type General Purpose
Context 4K
Reasoning 6/10
Speed 4/10
Fine-tuning Streaming
Status

Available open-weight model; not formally deprecated or retired

View model →

Local inference, open-model research, benchmarking, continued pretraining, and task-specific fine-tuning.

Type General Purpose
Context 4K
Reasoning 5/10
Speed 7/10
Fine-tuning
Status

Available as a downloadable open-weight model; no official first-party hosted API identified.

Input No official Ai2 hosted API price; self-hosted model weights are downloadable.
Output No official Ai2 hosted API price; self-hosted model weights are downloadable.
View model →

Fully open language-model research, local inference, reproducible training experiments, evaluation, and downstream fine-tuning

Type General Purpose
Context 4K
Reasoning 6/10
Speed 4/10
Fine-tuning Streaming
Status

Current downloadable open-weight model; no first-party hosted API deployment verified

Input No official hosted API price; downloadable weights are available
Output No official hosted API price; downloadable weights are available
View model →

Continued pretraining, supervised fine-tuning, reinforcement-learning research, language-model research, custom deployment, and applications requiring a transparent open training pipeline.

Type General Purpose
Context 66K
Reasoning 5/10
Speed 7/10
Fine-tuning Streaming
Status

Current and downloadable open-weight model

View model →

Open, self-hosted chat assistants, instruction following, local coding help, tool-calling agents, and model research

Type General Purpose
Context 66K
Reasoning 7/10
Speed 7/10
Tool use Fine-tuning Streaming
Status

Current open-weight model; available for download and self-hosted inference

View model →

Open research, mathematical reasoning, code generation, logic, long-context experimentation, and local self-hosted inference

Type Reasoning
Context 66K
Reasoning 9/10
Speed 6/10
Fine-tuning Streaming
Status

Active open-weight model

View model →

Open-weight research, continued pretraining, fine-tuning, programming, mathematics, reading comprehension, and long-context language-model experiments

Type General Purpose
Context 66K
Reasoning 6/10
Speed 4/10
Fine-tuning Streaming
Status

Current; open-weight and downloadable

Input No official hosted API price; downloadable weights
Output No official hosted API price; downloadable weights
View model →

Open-weight chat, tool-using assistants, multi-turn dialogue, synthetic data generation, self-hosting, and model research

Type General Purpose
Context 66K
Reasoning 7/10
Speed 5/10
Tool use Fine-tuning Streaming
Status

Available open-weight checkpoint; superseded by Olmo 3.1 32B Instruct for newer instruction-tuned deployments

Input No official first-party hosted API price; downloadable open weights
Output No official first-party hosted API price; downloadable open weights
View model →

Open-model reasoning research, mathematics, coding, long-context analysis, and self-hosted deployment

Type Reasoning
Context 66K
Reasoning 8/10
Speed 5/10
Fine-tuning Streaming
Status

Available as an open-weight model; superseded by the newer Olmo 3.1 32B Think model

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; deployment and inference costs are user-managed
View model →

Math-focused RLVR research, reinforcement-learning experiments, verifiable reward design, open-model evaluation, and further fine-tuning

Type Reasoning
Context 66K
Reasoning 7/10
Speed 7/10
Fine-tuning
Status

Current open-weight research checkpoint

View model →

Mathematical reasoning, coding, complex multi-step tasks, open-model research, local deployment, and reinforcement-learning research

Type Reasoning
Context 66K
Reasoning 8/10
Speed 5/10
Fine-tuning Streaming
Status

Current open-weight model

View model →

Open coding-model research, reinforcement learning from verifiable rewards, code-generation experiments, automated evaluation, and fine-tuning

Type Coding
Context 66K
Reasoning 5/10
Speed 7/10
Fine-tuning
Status

Current, open-weight, experimental RL-Zero coding checkpoint

View model →

Self-hosted chat assistants, instruction following, tool-use agents, synthetic data, and open-model research

Type General Purpose
Context 66K
Reasoning 7/10
Speed 4/10
Tool use Fine-tuning
Status

Current; open-weight and downloadable

View model →

English audio transcription, captions, meetings, lectures, podcasts, accessibility applications, and open ASR research

Type Other
Reasoning 1/10
Speed 8/10
Audio input Streaming
Status

Current open-weight model

View model →

Open English speech-to-text transcription, captioning, meetings, lectures, calls, podcasts, and local ASR research

Type Other
Reasoning 1/10
Speed 6/10
Audio input Streaming
Status

Available open-weight model

Input No official hosted API pricing; released for open-weight deployment
Output No official hosted API pricing; released for open-weight deployment
View model →

Open, self-hosted English transcription for meetings, lectures, calls, podcasts, accessibility, and speech-recognition research

Type Other
Reasoning 1/10
Speed 5/10
Audio input Streaming
Status

Current open-weight model; downloadable and usable for local inference

Input No official hosted API pricing; self-hosted open weights
Output No official hosted API pricing; self-hosted open weights
View model →

English short- and long-form speech transcription, meeting and call transcription, lecture captioning, podcast processing, broadcast transcription, and speech analytics

Type Other
Reasoning 1/10
Speed 6/10
Audio input
Status

Current open-weight model

Input No official hosted API price; self-hosted checkpoint
Output No official hosted API price; self-hosted checkpoint
View model →

Open English speech transcription, meeting and podcast transcription, captioning, timestamped audio indexing, and ASR research

Type Other
Reasoning 1/10
Speed 7/10
Audio input
Status

Current open-weight model; available for download and self-managed inference

Input No official hosted API pricing; open weights available for self-hosted use
Output No official hosted API pricing; output is transcribed text generated during local inference
View model →

Efficient self-hosted English transcription, speech research, edge-oriented experiments, and applications where a small open ASR checkpoint is preferred.

Type Lightweight
Reasoning 1/10
Speed 8/10
Audio input
Status

Available open-weight checkpoint

Input No official hosted API price; self-hosted checkpoint
Output No official hosted API price; self-hosted checkpoint
View model →

Open-model research, efficient local text generation, language-model evaluation, and fine-tuning

Type Lightweight
Context 4K
Reasoning 4/10
Speed 8/10
Fine-tuning Streaming
Status

Available open-weight release

View model →

Efficient satellite-image and Earth-observation embeddings, remote-sensing research, and downstream classification or segmentation

Type Multimodal
Reasoning 1/10
Speed 8/10
Multimodal Image input Fine-tuning
Status

Current; downloadable open-weight research model

View model →
Allen Institute for Artificial Intelligence (Ai2) logo
OlmoEarth v1.2

OlmoEarth-v1_2-Base

Satellite-image and Earth-observation embeddings, remote-sensing representation learning, geospatial classification, segmentation, and downstream fine-tuning

Type Multimodal
Reasoning 1/10
Speed 7/10
Multimodal Image input Fine-tuning
Status

Current open-weight model

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; produces feature representations and embeddings
View model →
Allen Institute for Artificial Intelligence (Ai2) logo
OlmoEarth v1.2

OlmoEarth-v1_2-Nano

Efficient Earth observation embeddings, satellite image and time-series representation learning, geospatial classification, segmentation, and large-scale remote-sensing inference

Type Multimodal
Reasoning 2/10
Speed 9/10
Multimodal Image input Fine-tuning
Status

Current open-weight model

View model →
Allen Institute for Artificial Intelligence (Ai2) logo
OlmoEarth v1.2

OlmoEarth-v1_2-Small

Satellite-image embeddings, remote-sensing representation learning, geospatial segmentation, land-cover analysis, and Earth observation research

Type Multimodal
Reasoning 1/10
Speed 7/10
Multimodal Image input Fine-tuning
Status

Current open-weight Earth observation foundation model

Input No official public hosted API price; downloadable weights are available
Output No official public hosted API price; model output is primarily embeddings and task representations
View model →
Allen Institute for Artificial Intelligence (Ai2) logo
Olmo Hybrid

Olmo Hybrid 7B

Open-weight language-model research, long-context text generation, local deployment, continued pretraining, and fine-tuning

Type General Purpose
Context 66K
Reasoning 6/10
Speed 8/10
Fine-tuning Streaming
Status

Current open-weight model

View model →

Self-hosted instruction following, open post-training research, mathematics, coding, evaluation and domain adaptation

Type General Purpose
Context 131K
Reasoning 7/10
Speed 7/10
Fine-tuning
Status

Available for download; superseded by Llama-3.1-Tulu-3.1-8B

View model →

Self-hosted instruction following, reasoning, mathematics, coding, open-model research, and reproducible post-training experiments

Type General Purpose
Context 131K
Reasoning 8/10
Speed 3/10
Fine-tuning
Status

Available open-weight model

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; downloadable weights
View model →

Large-scale research, instruction-following evaluation, open-weight post-training research, mathematical reasoning, coding benchmarks, and self-hosted experimentation.

Type General Purpose
Context 8K
Reasoning 8/10
Speed 2/10
Fine-tuning Streaming
Status

Available as an open-weight research and educational model; no hosted inference provider deployment is currently listed on its Hugging Face model page.

Input No official first-party hosted API price; self-hosted open weights
Output No official first-party hosted API price; self-hosted open weights
View model →
Allen Institute for Artificial Intelligence (Ai2) logo
Unified-IO

Unified-IO

Multimodal research, vision-language experiments, image generation, visual question answering, dense computer-vision tasks, and academic benchmarking

Type Multimodal
Reasoning 4/10
Speed 3/10
Multimodal Image input Media output
Status

Open-weight research release; publicly available inference code and checkpoints; legacy relative to Ai2's newer multimodal models

View model →
Allen Institute for Artificial Intelligence (Ai2) logo
Unified-IO

Unified-IO 2

Multimodal research, image understanding and generation, audio and video understanding, spatial prediction, embodied AI and robotic-manipulation experiments, and self-hosted academic prototyping.

Type Multimodal
Reasoning 5/10
Speed 2/10
Multimodal Image input Audio input
Status

Open-weight research release; publicly accessible checkpoints and source code; no official hosted inference API or commercial token pricing identified.

Input No official hosted API pricing; self-hosted research checkpoints
Output No official hosted API pricing; self-hosted research checkpoints
View model →
Allen Institute for Artificial Intelligence (Ai2) logo
WildDet3D

WildDet3D

Open-vocabulary monocular 3D detection, spatial perception, robotics research, augmented reality, and lifting 2D prompts into metric 3D boxes

Type Other
Reasoning 2/10
Speed 2/10
Multimodal Image input Fine-tuning
Status

Current open-weight research model

Input No official hosted API pricing; downloadable checkpoint
Output No official hosted API pricing; downloadable checkpoint
View model →