What OlmoEarth-v1_2-Nano is
OlmoEarth-v1_2-Nano is an open-weight Earth observation foundation model provided by the Allen Institute for AI, also known as Ai2. It is designed for remote-sensing workflows: applications that analyze images and measurements collected by satellites or other Earth observation systems.
The model is not a chatbot and does not generate prose. Its main job is to learn a useful numerical representation of geospatial data. That representation can then be supplied to a classifier, segmentation system, similarity search tool, or other downstream task. In practical terms, the model can help turn raw satellite observations into features that are easier for a specialized application to use.
OlmoEarth-v1_2-Nano is distributed through Ai2’s Hugging Face organization under the model name OlmoEarth-v1_2-Nano. It is identified as an image feature-extraction model and belongs to the OlmoEarth v1.2 family. Within that family, Nano is the compact configuration intended for lower-cost processing and efficient deployment.
Primary purpose and position in the OlmoEarth family
The model is intended for representation learning across satellite images and image time series. Representation learning means that the model converts complex inputs into embeddings that preserve information useful for later analysis. A user might use those embeddings to classify land cover, identify crop types, map regions, compare locations, or detect changes over time.
Compared with larger Earth observation models, the Nano configuration emphasizes compactness. Its encoder contains approximately 1.7 million parameters and its decoder contains approximately 800,000 parameters. That small size can reduce memory requirements and make it more practical for experimentation, high-throughput processing, or deployments where GPU resources are limited.
The compact design involves a trade-off. Nano is useful when processing efficiency matters, but it should not be treated as a universal replacement for larger models or a complete geospatial application. Most real deployments still need a task-specific classifier, decoder, prediction head, or fine-tuning procedure after feature extraction.
Supported inputs and modalities
OlmoEarth-v1_2-Nano is multimodal in the Earth observation sense: it can work with several types of geospatial data rather than a single ordinary photograph. Its documented inputs and training modalities include:
- Sentinel-2 L2A multispectral imagery
- Sentinel-1 radar imagery
- Landsat imagery
- WorldCover land-cover data
- SRTM elevation data
- OpenStreetMap raster data
- WRI Canopy Height Map data
- USDA Cropland Data Layer data
- WorldCereal data
Sentinel-1 supplies radar information, while Sentinel-2 and Landsat provide multispectral observations. The additional map and environmental layers can act as supervisory signals during pretraining or provide complementary geographic information. The exact preparation process matters: users need to provide data in the modality-specific formats, band orders, and normalization scheme expected by the model.
The configuration supports sequences of up to 12 timesteps. This is a limit on the supported Earth observation sequence, not a language-model context window. It should therefore be understood as a temporal input limit for image observations rather than a number of text tokens. The supplied specifications do not identify a conventional token context length or a maximum generated-output length.
How the model works
OlmoEarth-v1_2-Nano uses a ViT-Nano vision transformer in an encoder-decoder arrangement. A vision transformer divides image information into tokens and processes relationships among those tokens. In this model, the tokens represent portions of geospatial observations, including information from different bands and time steps.
Pretraining uses masked image modeling in latent token space. In simplified terms, portions of the input are hidden and the model learns to reconstruct or represent the missing information. The training approach combines multimodal and temporal inputs with reconstruction and contrastive objectives. Contrastive learning encourages representations of related observations to be meaningfully associated while distinguishing less-related examples.
Version 1.2 introduces several documented changes intended to improve efficiency and representation quality. These include one band set per modality, random band dropout for Sentinel-2 and Landsat, a nonlinear projection layer, updated temporal masking, revised contrastive-loss handling, and rotary positional encodings. Ai2’s technical material also describes the changes as addressing striping artifacts observed in earlier embeddings.
Outputs and downstream tasks
The primary output is a learned geospatial representation or embedding. An embedding is a numerical summary that a downstream system can use without requiring the original model to directly produce a final human-readable answer.
Potential uses supported by the model’s design include:
- Land-cover classification
- Crop-type classification
- Semantic segmentation and geospatial mapping
- Environmental monitoring
- Change detection across image dates
- Remote-sensing retrieval and similarity search
- Clustering satellite scenes or geographic regions
- Large-area feature extraction for later statistical analysis
OlmoEarth-v1_2-Nano does not natively output text, audio, video, or generated images. Its output is best understood as embeddings and reconstruction-related representations. Producing a map, label, report, or other end-user result requires additional software and usually a task-specific model component.
Strengths and trade-offs
Compact inference
The strongest practical distinction is the model’s small size. With approximately 1.7 million encoder parameters and 800,000 decoder parameters, it is positioned for lower-cost inference than larger foundation models. That can be valuable when processing many scenes, testing ideas on modest hardware, or deploying a model close to a data source.
Multisource and temporal analysis
The model is designed around more than one satellite source and can handle image sequences of up to 12 timesteps. This makes it more suitable for observing seasonal or environmental change than a system trained only on isolated single-date images, provided the input data is prepared correctly.
Open model access
Ai2 publishes the checkpoint through Hugging Face and provides associated training and inference resources. This gives technical users access to downloaded weights and research code rather than requiring a proprietary hosted endpoint. The model is distributed under Ai2’s OlmoEarth Artifact License, and users should review that license and Ai2’s Responsible Use Guidelines before using it in a commercial or other restricted setting.
Important limitations
Nano is not a general-purpose AI assistant. It has no documented text-generation, coding, web-search, function-calling, or tool-use capability. Its reasoning and coding scores in a model catalog should not be interpreted as benchmark results or provider claims; those capabilities are simply outside its intended role.
The model also does not remove the need for geospatial preprocessing. Differences in sensor calibration, geographic coverage, cloud conditions, missing bands, coordinate systems, temporal alignment, and normalization can affect downstream performance. The supplied research does not establish a universal accuracy score for every classification or segmentation task, so results should be validated on the target region and application.
Pricing and deployment
No official per-token price, hosted inference price, or recurring subscription price is specified for OlmoEarth-v1_2-Nano. The primary access model is downloading the open checkpoint and running it on user-managed infrastructure, although Ai2’s broader platform and project availability may change over time. Infrastructure costs therefore depend on the user’s hardware, hosting provider, processing volume, and preprocessing pipeline.
Because this is not a language model endpoint, common API metrics such as input tokens, output tokens, streaming, and JSON mode do not apply. The relevant deployment considerations are image and time-series throughput, memory use, storage, data-transfer costs, and the compute required by any downstream prediction head.
When to choose OlmoEarth-v1_2-Nano
Choose OlmoEarth-v1_2-Nano when you need an efficient, downloadable foundation model for satellite or geospatial representation learning. It is a reasonable starting point for teams that want to process large image collections, experiment with Sentinel or Landsat data, build a classifier on top of pretrained features, or reduce the compute cost of an Earth observation pipeline.
Its compact size is especially relevant when throughput and operating cost matter more than using the largest available representation model. It may also be appropriate for research groups that value open weights, published training resources, and the ability to adapt the model rather than relying on a provider-managed black-box service.
Another option may be more appropriate when the task requires conversational explanations, code generation, document analysis, speech processing, image generation, or an agent that calls external tools. A larger or more specialized remote-sensing model may also be preferable when maximum task accuracy is more important than compact inference, but the supplied research does not provide a direct benchmark comparison with named alternatives.
Bottom line
OlmoEarth-v1_2-Nano is a specialized and compact Earth observation encoder-decoder, not a general-purpose generative model. Its value comes from converting multisource, multitemporal satellite data into reusable embeddings at relatively low model size. For geospatial teams that can manage preprocessing and add their own downstream task components, it offers an efficient open-weight foundation for classification, mapping, monitoring, and large-scale remote-sensing analysis.

