OLMoASR

OLMoASR-small.en

by Allen Institute for Artificial Intelligence (Ai2) · Current open-weight model; available for download and self-managed inference

An open 244-million-parameter English speech recognition model from Ai2 that transcribes short- and long-form audio, provides sentence-level timestamps, and offers a middle ground between smaller and larger OLMoASR checkpoints.

Text Reasoning Coding
OLMoASR-small.en is an open-weight English speech-to-text model developed by the Allen Institute for AI (Ai2). It sits between the smaller OLMoASR-tiny.en and OLMoASR-base.en checkpoints and the larger medium and large variants, offering a practical balance between transcription accuracy and deployment cost. The model is intended for local or self-managed inference rather than a metered first-party cloud API.
Outputs

What OLMoASR-small.en can produce

Text
Inputs

What it can understand

Audio
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family OLMoASR
Model type Other
Release date 2025-08-28
Status Current open-weight model; available for download and self-managed inference
Knowledge cutoff notes

A model knowledge cutoff is not specified because this is an automatic speech recognition checkpoint rather than a general-purpose knowledge model. Its training data consists of English audio-text pairs collected and curated for speech recognition.

Model notes

OLMoASR-small.en is the 244-million-parameter member of Ai2's OLMoASR family. It uses a Transformer encoder-decoder architecture with an audio encoder and language decoder, is trained on English-only data, and produces sentence-level timestamps. Ai2 reports average word error rates of 13.8% on short-form evaluation sets and 11.5% on long-form evaluation sets. The model is distributed through the allenai/OLMoASR Hugging Face repository under Apache 2.0. The official usage path is Python-based through the olmoasr package. No official first-party hosted inference price, batch API, web-search capability, structured-output mode, or streaming interface is documented for this exact checkpoint. The editorial scores are comparative estimates for an automatic speech recognition model, not vendor ratings.

Cost

Model pricing

Input No official hosted API pricing; open weights available for self-hosted use
Output No official hosted API pricing; output is transcribed text generated during local inference
Model guide

OLMoASR-small.en: Open 244M-Parameter English Speech Recognition

OLMoASR-small.en is a 244-million-parameter, English-only automatic speech recognition model from the Allen Institute for AI. It converts speech into text, supports short- and long-form audio, and produces sentence-level timestamps for applications such as meeting transcription, captioning, podcast indexing, and speech-recognition research.

What is OLMoASR-small.en?

OLMoASR-small.en is an automatic speech recognition (ASR) model. ASR systems analyze recorded or live speech and produce written text. In this case, the model is designed specifically for English audio and returns transcription text with sentence-level timestamps.

The checkpoint contains approximately 244 million parameters. Parameters are the learned numerical values that allow a model to recognize speech patterns, words, accents, and audio conditions. OLMoASR-small.en uses a Transformer-based encoder-decoder architecture: an audio encoder interprets the sound signal, while a language decoder generates the corresponding written transcription.

Ai2 provides OLMoASR-small.en as part of the open OLMoASR family. The family includes tiny, base, small, medium, large, and large-v2 variants. The small model is therefore not the lowest-cost or smallest option, but it is also not the largest accuracy-oriented checkpoint in the lineup.

Primary purpose and supported inputs

The model's primary purpose is English speech transcription. It can be used with both short-form and long-form recordings, including meetings, telephone conversations, lectures, interviews, broadcasts, audiobooks, and podcasts.

Its supported input is English audio. The supplied documentation does not identify it as a multilingual model, image-capable model, video-understanding model, or general-purpose text-generation model. It produces text and timestamp metadata rather than synthesized speech, images, or video.

  • Input: English audio
  • Output: Transcribed English text with sentence-level timestamps
  • Audio understanding: Yes, for automatic speech recognition
  • Image, video, and text-generation input: Not documented for this checkpoint
  • Speech synthesis: Not supported

Sentence-level timestamps are useful when the transcription must remain connected to the source recording. For example, a captioning tool can use them to align text with playback, while a podcast search system can use them to direct a user to the relevant point in an episode.

Reported performance

Ai2 reports an average word error rate (WER) of 13.8% across its short-form evaluation sets and 11.5% across its long-form evaluation sets. Word error rate measures the proportion of recognized words that differ from the reference transcription; lower values indicate fewer word-level errors. These figures are provider-reported evaluation results, not a guarantee of performance on every recording.

In the reported benchmark suite, OLMoASR-small.en achieved the following WER results:

Evaluation setReported WER
LibriSpeech test-other7.0%
TED-LIUM34.2%
Switchboard13.2%
Earnings-22 long-form14.0%

Ai2 says that OLMoASR-small.en roughly matches Whisper-small.en on its short-form and long-form word error rate comparisons. Actual results can vary with background noise, microphone quality, speaker accents, overlapping speech, recording conditions, and the subject matter being discussed. Benchmark scores should therefore be treated as a useful reference point rather than a substitute for testing representative recordings.

Training and open release

OLMoASR models were trained from scratch using weakly supervised audio-text data collected from the public internet. Ai2 describes OLMoASR-Pool as a three-million-hour pool that was filtered into OLMoASR-Mix, a curated collection of approximately one million hours.

The open project includes model weights, training and data-processing code, evaluation code, and associated datasets. The small checkpoint is distributed through Ai2's OLMoASR repository on Hugging Face. The model card identifies the release under the Apache 2.0 license, while also noting that the checkpoint was trained on English-only data.

This openness is relevant to organizations that need to inspect, reproduce, customize, or run a speech model without depending entirely on a proprietary transcription service. It also shifts more responsibility to the user: deployment, infrastructure, monitoring, data handling, and compliance are not provided as a single managed service by the model itself.

Deployment and API availability

OLMoASR-small.en is intended for local or self-managed inference. Ai2 documents Python-based use through the olmoasr package, where a user loads the model and invokes its transcription functionality on an audio file.

No official first-party hosted inference price is specified for this exact checkpoint. There is also no documented managed API, batch API, streaming interface, or hosted per-minute billing model in the supplied information. Consequently, the practical cost is determined by the infrastructure used to run it, including compute, storage, operations, and engineering time.

The 244-million-parameter size gives the model a larger deployment footprint than the 39-million-parameter OLMoASR-tiny.en and 74-million-parameter OLMoASR-base.en variants. It will generally require more memory and compute than those smaller checkpoints. In exchange, Ai2 reports lower word error rates than the smaller variants. The supplied documentation does not specify a context window, maximum output-token limit, minimum hardware configuration, or a fixed real-time factor, so those values should not be assumed.

Capabilities that are not part of the model

OLMoASR-small.en is an ASR checkpoint rather than a general-purpose assistant or language model. Its role is to convert audio into text, not to reason over a conversation, write software, browse the web, call external tools, or produce structured application actions.

  • Reasoning: No general reasoning capability is documented. Any analysis of the transcript must be performed by a separate model or application.
  • Coding: No code-generation capability is documented.
  • Tool and function calling: Not supported or documented.
  • Web search: Not supported.
  • Structured output: No structured-output mode is documented, although the transcription includes sentence-level timestamp information.
  • Streaming: No streaming interface is documented for this exact checkpoint.

A transcription workflow can still be combined with other software. For example, an application could pass OLMoASR-small.en's transcript to a separate language model for summarization or extraction. That would be a multi-component system, however, and should not be attributed to OLMoASR-small.en itself.

Strengths and limitations

Main strengths

  • Open deployment: The weights and supporting project materials are available for self-managed use under the stated Apache 2.0 release.
  • Useful middle position: The small checkpoint provides a compromise between the lower resource requirements of tiny and base models and the higher resource requirements of medium and large models.
  • Long-form support: It is designed for extended recordings as well as short clips.
  • Timestamped results: Sentence-level timestamps support captions, search, indexing, and media navigation.
  • Research transparency: Ai2 releases training, processing, evaluation, and model-related artifacts rather than offering only a closed transcription endpoint.

Main limitations

  • English only: It is not a suitable primary choice for multilingual transcription.
  • Self-managed operations: Users must arrange hosting and inference instead of calling a documented first-party paid API.
  • Resource requirements: The model is larger and more demanding than the tiny and base checkpoints.
  • Variable real-world accuracy: Noise, overlapping speakers, accents, poor microphones, and unfamiliar domains can reduce transcription quality.
  • Narrow task scope: It does not synthesize speech, generate images or video, reason over content, or provide application tools.
  • Privacy responsibilities: Organizations must address consent, copyright, confidentiality, and data-protection obligations when processing recordings.

When to choose OLMoASR-small.en

Choose OLMoASR-small.en when you need an open English transcription model and can operate it locally or on infrastructure you control. It is a reasonable fit for meeting archives, podcast and lecture transcription, timestamped media search, captioning pipelines, call-recording analysis, and ASR experimentation where inspectable model artifacts matter.

Its middle position in the OLMoASR family is particularly useful when the smallest checkpoints do not provide enough accuracy, but the compute requirements of the medium or large variants are difficult to justify. The reported results also make it a relevant candidate for users comparing open speech recognition with similarly sized proprietary or open alternatives.

Another option may be more appropriate in several situations. Select a smaller OLMoASR variant when lower memory use, faster inference, or a smaller deployment footprint is more important than the reported accuracy trade-off. Consider a larger OLMoASR variant when accuracy is the priority and additional compute is available. Choose a managed transcription service when you need a hosted endpoint, predictable operational support, or usage-based billing rather than self-managed inference. Use a multilingual ASR system when the recordings contain languages beyond English.

Pricing and practical cost

There is no official hosted API price supplied for OLMoASR-small.en. The open checkpoint can be downloaded for self-managed use, but “open” does not mean cost-free in production: users may still pay for GPUs or other compute, storage, networking, monitoring, and engineering support.

Because no fixed hosted price, context limit, maximum output limit, or streaming guarantee is documented for this exact model, teams should benchmark it on representative audio before estimating throughput or total cost. A useful evaluation should include the recording lengths, accents, noise levels, speaker overlap, and vocabulary found in the intended application.

Bottom line

OLMoASR-small.en is a focused, open English speech-recognition model for users who want local control, timestamped transcription, and a balance between model size and reported accuracy. Its value is strongest in self-managed transcription and research workflows. It is not a hosted conversational assistant or a general-purpose multimodal model, and its English-only scope, infrastructure requirements, and lack of a documented first-party API should be considered before deployment.


Answers to Frequently Asked Questions

How large is OLMoASR-small.en and what license does it use?
OLMoASR-small.en contains approximately 244 million parameters. The model card identifies the release under the Apache 2.0 license, and the open project includes model weights, training and data-processing code, evaluation code, and associated datasets.
Is OLMoASR-small.en available through a hosted API?
The supplied documentation does not specify an official hosted inference price, managed API, batch API, streaming interface, or per-minute billing model for OLMoASR-small.en. It is intended for local or self-managed inference through the Python-based olmoasr package, so users must provide the required compute, storage, operations, and monitoring.
Does OLMoASR-small.en support languages other than English?
No. OLMoASR-small.en is designed for English audio and was trained on English-only data. A multilingual ASR model should be used when recordings contain other languages.
How accurate is OLMoASR-small.en?
Ai2 reports an average word error rate of 13.8% across short-form evaluation sets and 11.5% across long-form evaluation sets. Reported results include 7.0% on LibriSpeech test-other, 4.2% on TED-LIUM3, 13.2% on Switchboard, and 14.0% on Earnings-22 long-form. Actual accuracy varies with noise, accents, microphone quality, overlapping speech, and subject matter.
What is OLMoASR-small.en used for?
OLMoASR-small.en is an automatic speech recognition model for transcribing English audio into written text with sentence-level timestamps. It can be used for meetings, interviews, lectures, podcasts, broadcasts, telephone recordings, audiobooks, captioning, and timestamped media search.


Sources 4
Provider

About Allen Institute for Artificial Intelligence (Ai2)