OLMo 2

OLMo-2-1124-13B

by Allen Institute for Artificial Intelligence (Ai2) · Current downloadable open-weight model; no first-party hosted API deployment verified

Ai2’s OLMo-2-1124-13B is a fully open 13-billion-parameter base language model for research, local inference, and fine-tuning. It supports text input and output, has a 4,096-token context length, and was released with weights, training data, code, recipes, checkpoints, and evaluation resources. No official hosted API pricing or native multimodal and tool-use capabilities are verified.

Text Reasoning Coding
OLMo-2-1124-13B is the base 13B model in the Allen Institute for AI’s OLMo 2 family. It is designed less as a ready-made chat assistant and more as an inspectable foundation for researchers and developers who need downloadable weights, reproducible training materials, local deployment, and downstream adaptation. The model supports text input and output, uses a 4,096-token context length, and is available under the Apache 2.0 license. There is no verified first-party hosted API or official per-token pricing for this checkpoint.
Outputs

What OLMo-2-1124-13B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

6/10 Reasoning
5/10 Coding
4/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family OLMo 2
Model type General Purpose
Context window 4K tokens
Knowledge cutoff December 2023
Release date November 26, 2024
Status Current downloadable open-weight model; no first-party hosted API deployment verified
Knowledge cutoff notes

The official model card states a date cutoff of December 2023. This is the training-data cutoff for the base model and is not changed by external retrieval, prompting, or downstream integrations.

Model notes

Canonical Hugging Face identifier is allenai/OLMo-2-1124-13B. This is the base pretrained model, not the separately released OLMo-2-1124-13B-Instruct variant. Ai2 released the model with weights, training data, code, recipes, evaluation resources, and intermediate checkpoints. The model card documents English language support, approximately 13 billion parameters, 40 layers, a hidden size of 5,120, 40 attention heads, a 4,096-token context length, and a December 2023 data cutoff. No official per-token hosted API pricing or first-party web-search capability was verified. Streaming is available through compatible local generation frameworks rather than a provider-specific hosted API. Editorial scores are comparative estimates, not vendor specifications.

Cost

Model pricing

Input No official hosted API price; downloadable weights are available
Output No official hosted API price; downloadable weights are available
Model guide

OLMo-2-1124-13B: A Fully Open 13B Model for Local Research

OLMo-2-1124-13B is Ai2’s fully open, 13-billion-parameter base language model, released with weights, training data, code, recipes, checkpoints, and evaluation resources for transparent research, local inference, and fine-tuning.

What is OLMo-2-1124-13B?

OLMo-2-1124-13B is a 13-billion-parameter autoregressive language model from the Allen Institute for AI, also known as Ai2. An autoregressive model generates text by predicting the next token based on the text that comes before it. In practical terms, this checkpoint can be used for text completion, evaluation, continued training, and fine-tuning into more specialized applications.

The “base” designation is important. OLMo-2-1124-13B is pretrained rather than instruction-tuned, so it is not intended to behave like a polished conversational assistant out of the box. It may complete prompts effectively, but reliable task following, assistant-style dialogue, safety behavior, and structured workflows generally require suitable prompting, additional fine-tuning, or an application layer around the model.

Ai2 positions OLMo 2 as a fully open language-model project. The release is broader than simply publishing downloadable weights: it includes training data mixtures, source code, training recipes, evaluation resources, and intermediate checkpoints. That makes this model particularly relevant to researchers studying data composition, optimization, interpretability, reproducibility, and model behavior.

Where it fits in Ai2’s lineup

OLMo-2-1124-13B is part of Ai2’s OLMo 2 family and is the pretrained 13B base variant. It should not be confused with the separately released OLMo-2-1124-13B-Instruct model, which is adapted for instruction-following use cases. The model’s role in the lineup is to provide an open, general-purpose text foundation that can be examined and adapted rather than a turnkey hosted assistant.

Its open-release approach also distinguishes it from many commercial language models. Commercial models commonly emphasize managed access, rapid inference, support, and integrated tools, while OLMo-2-1124-13B emphasizes inspectability and control. Users can download the checkpoint, run it with compatible local software, examine the associated materials, and fine-tune it for a particular domain or task.

Architecture and training details

According to Ai2’s supplied model information, OLMo-2-1124-13B was trained in two major stages. The first stage used the OLMo-Mix-1124 data mixture and processed approximately 5 trillion tokens. A second stage used the Dolmino-Mix-1124 mixture, with multiple training runs whose resulting checkpoints were averaged into the released model.

The documented architecture contains 40 transformer layers, a hidden size of 5,120, and 40 attention heads. It has a 4,096-token context length. A context window is the amount of text the model can consider in a single request, including the prompt and generated continuation where the implementation counts both together. This is sufficient for many focused documents and experiments, but it is shorter than the long-context windows offered by some newer hosted systems.

Ai2’s technical materials describe rotary positional embeddings, RMSNorm-related improvements, QK normalization, and other training-stability techniques. These details matter most to users studying or modifying the training process; ordinary users can treat them as evidence that the release provides more architectural and training transparency than a conventional closed endpoint.

Capabilities and evaluation

As a base text model, OLMo-2-1124-13B is suited to language completion and downstream adaptation. Possible uses include experimenting with prompting, continued pretraining, domain fine-tuning, text generation, evaluation pipelines, and research into model behavior. It accepts text and produces text. The supplied specifications do not verify native image, audio, or video input, and the model does not directly generate images, audio, or video.

Ai2 reported an average score of 68.3 across its core evaluation table for this model. The reported table includes tasks such as ARC Challenge, HellaSwag, WinoGrande, MMLU, DROP, Natural Questions, AGIEval, GSM8K, MMLU-Pro, and TriviaQA. These are provider-reported evaluation results, not a guarantee of performance on a particular application. Results can also depend on prompting, decoding settings, benchmark version, and comparison models.

The editorial assessment supplied for this entry rates reasoning at 6 out of 10 and coding at 5 out of 10. These are comparative editorial scores rather than Ai2-published specifications. They suggest a useful general-purpose research foundation, but not a reason to assume that the model will match specialized coding or reasoning systems.

Deployment, access, and pricing

The canonical model identifier is allenai/OLMo-2-1124-13B on Hugging Face. The checkpoint can be loaded through the Hugging Face Transformers ecosystem, and Ai2 provides training and fine-tuning code through its OLMo repositories. Compatible quantization tools, including tools based on bitsandbytes, may reduce memory requirements, although a 13-billion-parameter model still requires substantial hardware resources compared with smaller language models.

This is a downloadable model rather than a documented pay-per-token service. No official hosted API price, monthly subscription price, or first-party inference endpoint was verified for this checkpoint. Therefore, the direct model price is best understood as no provider-listed hosted API charge, while the real operating cost depends on the hardware, hosting service, electricity, storage, and engineering work used to run it.

Local generation frameworks may support streaming output, but this should not be confused with a provider-specific streaming API. Similarly, no first-party web-search service, function-calling interface, tool-use layer, or official JSON mode is verified in the supplied specifications. Developers can build these features around the model, but the model itself should not be represented as providing them natively.

Modalities and practical limits

OLMo-2-1124-13B is text-only in both its verified input and output modalities. It does not natively process images, audio, or video. Users who need multimodal understanding should select a model specifically designed for those inputs rather than assuming that the broader Ai2 ecosystem’s multimodal projects apply to this checkpoint.

The documented context length is 4,096 tokens. A maximum output-token limit separate from that context length was not verified. In practice, the available generation space is constrained by the implementation and by the total context budget. The model’s knowledge cutoff is December 2023, so it cannot reliably answer questions about later events without an external retrieval system. It also has no built-in current-information access or verified web browsing.

Because this is an unfiltered base model, it can produce inaccurate, biased, unsafe, or unsuitable content. It may also respond unpredictably to conversational prompts that an instruction-tuned assistant would handle more consistently. Applications should validate outputs, apply appropriate safety controls, and independently verify factual claims, especially in high-impact settings.

Main strengths and trade-offs

  • Openness: The release includes weights and substantial training and evaluation materials, supporting reproducible research and inspection.
  • Adaptability: The base checkpoint can be fine-tuned or used for controlled experiments involving data, optimization, evaluation, and interpretability.
  • Local control: Users can deploy the model through compatible software instead of sending prompts to a mandatory first-party hosted service.
  • Research transparency: Intermediate checkpoints, recipes, and data-mixture information are useful for studying how models are trained.
  • Operational cost: There is no verified hosted API fee, but local operation shifts costs to hardware, hosting, storage, and maintenance.
  • Usability limitations: The base model is not a ready-made chat assistant and has no verified native tools, web search, multimodal input, or official JSON-output mode.

These trade-offs make OLMo-2-1124-13B different from smaller, faster models optimized for inexpensive production inference and from managed commercial assistants optimized for convenience. Its value is strongest when transparency and customization matter more than a turnkey user experience.

Best use cases

This model is a good fit for researchers who need an openly available checkpoint and accompanying artifacts for reproducible experiments. It is also suitable for developers who want to fine-tune a general text model for a controlled domain, compare training methods, test evaluation procedures, or run inference locally under their own infrastructure.

For example, a research team could compare the behavior of different checkpoints, investigate how a domain-specific dataset changes performance, or build a text-generation prototype without depending on a proprietary API. An organization with appropriate hardware could also adapt the model for an internal text task, provided it reviews the license, data provenance, security requirements, and output risks.

When to choose OLMo-2-1124-13B

Choose OLMo-2-1124-13B when you specifically value downloadable weights, transparent training materials, local control, and the ability to fine-tune or inspect the model. It is especially compelling for open-model research, education, benchmarking, and experiments where the training process matters as much as the final generated text.

Choose another type of model when you need a polished conversational experience, guaranteed hosted availability, built-in web search, native function calling, current information, multimodal input, or a simple usage-based API. A smaller model may also be more appropriate when low latency and low hardware cost are the priority. Conversely, a specialized instruction-tuned or coding model may be preferable when dependable task following or software development performance is more important than access to the base training artifacts.

License and responsible use

OLMo 2 is released under the Apache 2.0 license. Ai2 also provides responsible-use guidance and model documentation. The license does not remove the need to review the model card, data provenance, downstream fine-tuning data, privacy obligations, and sector-specific requirements. Before deploying the model commercially or in a high-impact application, teams should test it on representative data and establish monitoring and human-review procedures.

Overall, OLMo-2-1124-13B is best understood as an open research foundation rather than a finished assistant. Its combination of a 13-billion-parameter text model, downloadable checkpoint, broad training artifacts, and Apache 2.0 licensing gives technically capable users meaningful control. That same design means users must supply the infrastructure, adaptation, evaluation, and safeguards that a managed assistant would ordinarily provide.


Answers to Frequently Asked Questions

What license does OLMo-2-1124-13B use, and is there an official API price?
OLMo-2-1124-13B is released under the Apache 2.0 license. It is a downloadable model rather than a documented pay-per-token service, and no official hosted API price or first-party inference endpoint was verified. Operating costs depend on hardware, hosting, electricity, storage, and maintenance.
What are the context length, modalities, and knowledge cutoff of OLMo-2-1124-13B?
OLMo-2-1124-13B has a documented context length of 4,096 tokens and is text-only for both input and output. It does not natively process images, audio, or video, and its knowledge cutoff is December 2023. It has no verified built-in web browsing or current-information access.
How can I run OLMo-2-1124-13B locally?
The canonical model identifier is allenai/OLMo-2-1124-13B on Hugging Face. It can be loaded through the Hugging Face Transformers ecosystem, while Ai2 provides related training and fine-tuning code through its OLMo repositories. Quantization tools may reduce memory requirements, but running a 13-billion-parameter model still requires substantial hardware.
What is OLMo-2-1124-13B?
OLMo-2-1124-13B is a 13-billion-parameter autoregressive language model from the Allen Institute for AI (Ai2). It is a pretrained base model intended for text completion, evaluation, continued training, local deployment, and fine-tuning rather than turnkey conversational use.
Is OLMo-2-1124-13B an instruction-tuned chat model?
No. OLMo-2-1124-13B is a base pretrained model, not an instruction-tuned assistant. Reliable task following, conversational behavior, safety controls, and structured workflows generally require prompting, fine-tuning, or an application layer.


Sources 6
Provider

About Allen Institute for Artificial Intelligence (Ai2)