What is OlmoEarth-v1_2-Small?
OlmoEarth-v1_2-Small is an Earth observation foundation model provided by the Allen Institute for Artificial Intelligence, commonly known as Ai2. A foundation model is a pretrained model that can provide a starting point for many specialized tasks instead of being trained from scratch for only one application.
For this model, the main output is a learned representation, often called an embedding. An embedding is a numerical description of an input image or image sequence that downstream software can use for classification, segmentation, retrieval, monitoring, or prediction. The model is therefore aimed at geospatial and remote-sensing workflows rather than conversation, text generation, or general-purpose reasoning.
The Small variant contains approximately 36 million parameters in its vision-transformer encoder. Ai2’s documentation also describes a separately documented decoder with approximately 7.4 million parameters. The model belongs to the OlmoEarth v1.2 family and is positioned as an efficient option for extracting features from Earth observation data.
What data can it process?
OlmoEarth-v1_2-Small is designed to work with several types of Earth observation input. The documented sources include Sentinel-1 radar imagery, Sentinel-2 optical imagery, and Landsat imagery. It can also work with image time series, allowing software to represent changes in an area across multiple observations rather than treating every image as an isolated snapshot.
The wider OlmoEarth training setup incorporates additional geographic and environmental layers, including OpenStreetMap, WorldCover, the USDA Cropland Data Layer, SRTM elevation data, the WRI Canopy Height Map, and WorldCereal. These layers provide information that satellite pixels alone may not capture, such as land-cover labels, elevation, roads, crop categories, or vegetation structure.
In practical terms, the model supports a form of multimodal input based on different remote-sensing and geospatial data sources. This should not be confused with a consumer multimodal assistant that accepts arbitrary photographs, documents, and spoken questions. The supplied specifications identify image input but do not identify text, audio, or video input for this checkpoint.
Architecture and the v1.2 changes
OlmoEarth-v1_2-Small uses a ViT-Small configuration. ViT, short for Vision Transformer, is a vision architecture that divides images into smaller regions and learns relationships between those regions. In this model, the architecture is adapted to remote-sensing data and image sequences rather than ordinary consumer photography.
The v1.2 release uses rotary positional encodings instead of absolute positional encodings. Positional encodings help a model understand where image patches occur in relation to one another. Ai2 describes the v1.2 change as improving efficiency and removing striping artifacts that appeared in embeddings from earlier versions. This is a provider-documented architectural change, not an independent benchmark result supplied with this page.
The Small configuration trades model size for lower resource requirements compared with a larger foundation model. The available research does not provide a complete hardware benchmark, latency table, or comparative accuracy result for this exact checkpoint, so deployment speed will depend on image dimensions, time-series length, batch size, hardware, preprocessing, and the downstream task.
What is the model used for?
The model’s central use case is representation learning: turning satellite and geospatial inputs into features that another model or analysis pipeline can use. Those features may be useful when a team has limited labeled data and wants to fine-tune a pretrained encoder rather than build a remote-sensing model from the beginning.
- Satellite embedding extraction: Generate numerical features for imagery, locations, or image time series.
- Land-cover classification: Support models that assign categories such as forest, cropland, water, or built-up land.
- Segmentation: Provide features for models that label individual pixels or regions.
- Agricultural analysis: Help downstream systems study cropland, seasonal patterns, or environmental conditions.
- Environmental monitoring: Support analysis of vegetation, land use, elevation-related patterns, and changes over time.
- Remote-sensing research: Provide an open starting point for experiments involving optical, radar, temporal, and map-derived data.
Fine-tuning is supported according to the supplied model record. In this context, fine-tuning means adapting the pretrained representation to a particular labeled dataset or task. The model is not itself a complete segmentation or classification application; users generally need to build or select the downstream task head, preprocessing pipeline, evaluation process, and deployment system.
Outputs, reasoning, coding, and tools
OlmoEarth-v1_2-Small primarily produces embeddings and learned image representations. It does not generate text, images, video, audio, music, or speech. It is consequently not suitable for asking questions in natural language and receiving a written answer, even though its input pipeline can combine several forms of Earth observation information.
Conventional language-model capabilities are not applicable. The supplied model record rates reasoning and coding capability at a minimal level for catalog purposes, but those values should not be interpreted as provider-published benchmark scores. This checkpoint is not a reasoning model or a code-generation model. It also has no documented tool-use, function-calling, web-search, streaming, batch-API, caching, or JSON-mode capability.
There is no published context window or maximum output-token limit for this model. Those fields are designed mainly for generative language models and do not map cleanly to an embedding model. Actual input limits will instead depend on the model implementation, supported image dimensions, number of temporal observations, available memory, preprocessing requirements, and the configuration used by the loading code.
Availability, license, and pricing
Ai2 publishes the model weights through Hugging Face under the OlmoEarth Artifact License. The official olmoearth_pretrain repository provides open-source code and model-loading utilities for embedding extraction and downstream fine-tuning. Ai2 also states that OlmoEarth models are deployed through the OlmoEarth Platform, although platform access and partner availability may differ from access to downloadable weights.
No official public hosted-inference price is specified for OlmoEarth-v1_2-Small. The model therefore has no verified per-token input or output price, monthly subscription price, or standard hosted API rate to report. Downloading and running the weights may still involve infrastructure costs such as GPU time, storage, data transfer, and engineering work. Those costs are separate from a provider-published model price and vary substantially by deployment.
Users should review the OlmoEarth Artifact License and Ai2’s responsible-use requirements before using the model commercially or in a production system. The availability of a public checkpoint does not by itself establish that every dataset, downstream use, or deployment scenario has identical permissions.
Strengths and limitations
The main strength of OlmoEarth-v1_2-Small is specialization. It is built for Earth observation and can represent information from radar, optical satellite imagery, temporal observations, and related geographic layers. That makes it more relevant to remote-sensing workflows than a general computer-vision or language model that was not trained around these data types.
Its approximately 36-million-parameter encoder also makes it a comparatively compact choice within a foundation-model workflow. A smaller model can be easier to store, fine-tune, and deploy than a larger alternative, although the supplied research does not establish a universal speed or accuracy advantage. The v1.2 positional-encoding changes are another stated benefit, particularly for efficiency and the removal of earlier striping artifacts.
The limitations are equally important. The model does not provide a finished application, conversational interface, generated report, or task-specific prediction without additional software. Results depend on sensor selection, temporal coverage, geographic distribution, preprocessing, labels, and fine-tuning. A model trained or adapted for one region may not perform equally well in another region or under different seasonal and atmospheric conditions.
There is also no conventional language-model knowledge cutoff, context window, maximum generated-token setting, or public hosted-API pricing schedule for this checkpoint. Those omissions are not missing chat features; they reflect the fact that OlmoEarth-v1_2-Small is an Earth observation representation model rather than a text-generation system.
When to choose OlmoEarth-v1_2-Small
Choose OlmoEarth-v1_2-Small when the project needs an open, remote-sensing-specific encoder for satellite imagery or image time series and the team is prepared to operate a model locally or integrate it into a research pipeline. It is particularly suitable when the desired result is an embedding for later classification, segmentation, retrieval, or geospatial prediction.
The Small configuration may be a practical starting point when deployment resources are limited or when rapid experimentation matters more than using the largest available representation model. Its open weights and loading utilities can also be useful for researchers who need inspectable artifacts and the ability to fine-tune the model rather than rely on a closed hosted service.
Another type of option may be more appropriate when the requirement is natural-language analysis, code generation, interactive question answering, image generation, speech, or a managed API with published service-level pricing. A task-specific remote-sensing model may also be preferable when a project requires a narrowly defined output, documented regional accuracy, or a production-ready segmentation and classification service. In those cases, OlmoEarth-v1_2-Small is best treated as a foundation component, not as the entire application.
Bottom line
OlmoEarth-v1_2-Small is a compact Ai2 Earth observation foundation model for converting satellite and geospatial observations into reusable features. Its value lies in its remote-sensing specialization, support for several Earth observation sources, open distribution, and suitability for downstream fine-tuning. It should be evaluated as an embedding and representation model, not as a chatbot or generative AI system. Teams choosing it should plan for their own preprocessing, downstream task design, infrastructure, evaluation, and license review.

