What is OlmoEarth-v1_2-Base?
OlmoEarth-v1_2-Base is an open-weight Earth-observation foundation model provided by the Allen Institute for AI, also known as Ai2. It is designed for remote-sensing workflows: analyzing images collected by satellites and combining spatial information with observations taken over time.
The model is not a conversational language model. Its main job is to extract useful numerical representations, commonly called embeddings, from satellite images or image time series. Those representations can then be passed to a classifier, segmentation system, regressor, or another task-specific application. In practical terms, the model supplies a reusable visual and geospatial starting point instead of directly producing a written answer.
On Hugging Face, the model is identified as an image feature-extraction model. It belongs to the OlmoEarth v1.2 family and is positioned as the Base variant, with a focus on reducing training and inference costs while retaining the performance characteristics of earlier OlmoEarth versions.
Architecture and model scale
OlmoEarth-v1_2-Base uses a ViT-Base architecture. ViT, or Vision Transformer, is a neural-network architecture that processes an image as a collection of visual patches and learns relationships between those patches. For OlmoEarth, that visual processing is adapted to Earth-observation data and temporal sequences rather than ordinary consumer photographs alone.
The model has approximately 114 million encoder parameters and a 30-million-parameter decoder. The encoder is the part most directly associated with producing representations from input data, while the decoder supports the model's broader pretraining design. These figures describe the model architecture and should not be confused with a context-window or maximum-output-token specification used for text-generation models.
According to the supplied OlmoEarth v1.2 technical report, the v1.2 changes reduce Base-model training requirements by approximately three times in GPU hours. The same research reports an approximately 2.9-times reduction in multiply-accumulate operations, or MACs, on Sentinel-2 tasks while maintaining overall performance. These are claims from the technical research rather than a guarantee of identical savings in every deployment or hardware environment.
Supported Earth-observation data and modalities
OlmoEarth-v1_2-Base is built for multimodal geospatial input. Its training data includes satellite imagery from Sentinel-1, Sentinel-2, and Landsat, along with derived geographic layers. The listed derived sources include OpenStreetMap raster data, ESA WorldCover, the USDA Cropland Data Layer, SRTM elevation data, the WRI Canopy Height Map, and WorldCereal.
Sentinel-1 contributes radar observations, while Sentinel-2 and Landsat provide optical satellite imagery. The model's intended use of these sources is not simply to inspect one picture at a time. It is also designed for image time series, allowing downstream systems to use changes across dates as part of their analysis.
This makes the model relevant to tasks such as land-cover mapping, crop and agricultural monitoring, ecosystem assessment, wildfire analysis, vegetation studies, and other applications where location, sensor type, and time matter together. The exact input preparation remains important: geographic coverage, sensor combination, temporal sampling, resolution, and preprocessing can all affect results.
What the model produces
The primary output is a learned feature representation or embedding. An embedding is a numerical description of the input that preserves patterns the model learned during pretraining. Applications can use these representations to compare locations, train a smaller prediction head, classify land cover, or support pixel-level and region-level analysis.
Ai2's documented workflows include embedding extraction and segmentation fine-tuning, as well as integration with rslearn. Fine-tuning means adapting the pretrained model to a particular dataset or task instead of training an entire Earth-observation system from the beginning.
OlmoEarth-v1_2-Base does not natively produce text, images, audio, or video. It is therefore better understood as a representation-learning model than as a generative AI assistant. It also does not expose a documented token-based chat interface, function-calling system, web-search tool, or general-purpose agent framework.
Pricing, access, and licensing
There is no official hosted API price supplied for OlmoEarth-v1_2-Base. The model is available as downloadable weights rather than as a conventional pay-per-token endpoint. As a result, there is no verified input price, output price, monthly subscription, context-window charge, or maximum-output-token allowance to report.
Users should distinguish the absence of hosted API pricing from zero operational cost. Running the model locally or on rented infrastructure can still require suitable hardware, storage, data preparation, and engineering work. Actual cost depends on the size of the imagery pipeline and the hardware used for inference or fine-tuning.
The downloadable weights use the OlmoEarth Artifact License. Anyone planning redistribution, commercial deployment, or use with sensitive geospatial data should review the license, model documentation, and Ai2's responsible-use guidance before deployment. The model is also associated with the OlmoEarth Platform, Ai2's hosted environment for Earth-observation workflows, but the supplied information does not provide a specific platform price or imply that every model feature is available through one unified hosted interface.
Capabilities and limitations
OlmoEarth-v1_2-Base's strongest capability is specialized Earth-observation representation learning. It is trained across several satellite and derived geospatial modalities, supports image time series, and can be adapted for downstream tasks. Its open-weight distribution can also be useful to research teams that need to inspect, fine-tune, or run a model outside a closed commercial API.
The model has important boundaries:
- It is not a text-generation model and cannot serve as a conversational assistant.
- It has no documented token context length or maximum output-token limit because its output is a feature representation rather than generated text.
- It does not natively generate images, audio, or video.
- It is not documented as supporting tools, function calls, streaming responses, web search, or structured text output.
- Results can vary with geographic region, sensor combination, temporal coverage, preprocessing, and downstream adaptation.
- It should be validated on representative local data before being used for operational, environmental, agricultural, or emergency decisions.
The model card and supplied research do not document a conventional text reasoning or coding capability. Any reasoning or coding score associated with the catalog record is an editorial database assessment, not a provider-published benchmark or claim. For the same reason, the model's relatively favorable cost and speed assessments should be treated as comparative metadata rather than guaranteed performance on a particular machine.
When to choose OlmoEarth-v1_2-Base
Choose OlmoEarth-v1_2-Base when the central problem involves satellite imagery, geospatial data, or observations collected over time. It is a suitable candidate for teams that want to:
- extract embeddings from Sentinel-1, Sentinel-2, Landsat, or related data;
- build land-cover or crop-monitoring classifiers;
- fine-tune a model for segmentation or other remote-sensing tasks;
- experiment with open Earth-observation foundation models;
- run and inspect downloadable weights rather than depend on a token-priced language-model API.
Its Base configuration may be especially appropriate when computational efficiency matters. The reported reductions in training GPU hours and MACs are useful positioning signals for teams balancing model quality against infrastructure cost, although actual savings should be measured on the team's own pipeline.
When another type of option may be better
A general-purpose vision-language model is a more appropriate choice when the application requires natural-language explanations, image question answering, or interaction with users. A language model is the better fit for chat, code generation, document drafting, tool use, and token-based APIs. A conventional computer-vision model may also be preferable when the task is narrow, the input format is fixed, or a smaller task-specific model is easier to operate.
OlmoEarth-v1_2-Base is also not automatically the best choice for every geospatial deployment. A commercial hosted remote-sensing service may reduce infrastructure and preprocessing work, while a model trained specifically for a particular region, sensor, or prediction target may perform better after local evaluation. The benefit of OlmoEarth is its open, reusable Earth-observation representation rather than a promise of universal accuracy.
Bottom line
OlmoEarth-v1_2-Base is a specialized, open-weight foundation model for turning satellite imagery and geospatial time series into useful representations. Its 114-million-parameter ViT-Base encoder, broad Earth-observation training sources, and support for downstream fine-tuning make it relevant to research and production prototyping in remote sensing. It should not be evaluated like a chat model: there is no text context window, token output limit, conversational reasoning interface, or official hosted API price. The main decision is whether an open Earth-observation embedding model matches the application's data and deployment needs.

