What is OLMoASR-small.en?
OLMoASR-small.en is an automatic speech recognition (ASR) model. ASR systems analyze recorded or live speech and produce written text. In this case, the model is designed specifically for English audio and returns transcription text with sentence-level timestamps.
The checkpoint contains approximately 244 million parameters. Parameters are the learned numerical values that allow a model to recognize speech patterns, words, accents, and audio conditions. OLMoASR-small.en uses a Transformer-based encoder-decoder architecture: an audio encoder interprets the sound signal, while a language decoder generates the corresponding written transcription.
Ai2 provides OLMoASR-small.en as part of the open OLMoASR family. The family includes tiny, base, small, medium, large, and large-v2 variants. The small model is therefore not the lowest-cost or smallest option, but it is also not the largest accuracy-oriented checkpoint in the lineup.
Primary purpose and supported inputs
The model's primary purpose is English speech transcription. It can be used with both short-form and long-form recordings, including meetings, telephone conversations, lectures, interviews, broadcasts, audiobooks, and podcasts.
Its supported input is English audio. The supplied documentation does not identify it as a multilingual model, image-capable model, video-understanding model, or general-purpose text-generation model. It produces text and timestamp metadata rather than synthesized speech, images, or video.
- Input: English audio
- Output: Transcribed English text with sentence-level timestamps
- Audio understanding: Yes, for automatic speech recognition
- Image, video, and text-generation input: Not documented for this checkpoint
- Speech synthesis: Not supported
Sentence-level timestamps are useful when the transcription must remain connected to the source recording. For example, a captioning tool can use them to align text with playback, while a podcast search system can use them to direct a user to the relevant point in an episode.
Reported performance
Ai2 reports an average word error rate (WER) of 13.8% across its short-form evaluation sets and 11.5% across its long-form evaluation sets. Word error rate measures the proportion of recognized words that differ from the reference transcription; lower values indicate fewer word-level errors. These figures are provider-reported evaluation results, not a guarantee of performance on every recording.
In the reported benchmark suite, OLMoASR-small.en achieved the following WER results:
| Evaluation set | Reported WER |
|---|---|
| LibriSpeech test-other | 7.0% |
| TED-LIUM3 | 4.2% |
| Switchboard | 13.2% |
| Earnings-22 long-form | 14.0% |
Ai2 says that OLMoASR-small.en roughly matches Whisper-small.en on its short-form and long-form word error rate comparisons. Actual results can vary with background noise, microphone quality, speaker accents, overlapping speech, recording conditions, and the subject matter being discussed. Benchmark scores should therefore be treated as a useful reference point rather than a substitute for testing representative recordings.
Training and open release
OLMoASR models were trained from scratch using weakly supervised audio-text data collected from the public internet. Ai2 describes OLMoASR-Pool as a three-million-hour pool that was filtered into OLMoASR-Mix, a curated collection of approximately one million hours.
The open project includes model weights, training and data-processing code, evaluation code, and associated datasets. The small checkpoint is distributed through Ai2's OLMoASR repository on Hugging Face. The model card identifies the release under the Apache 2.0 license, while also noting that the checkpoint was trained on English-only data.
This openness is relevant to organizations that need to inspect, reproduce, customize, or run a speech model without depending entirely on a proprietary transcription service. It also shifts more responsibility to the user: deployment, infrastructure, monitoring, data handling, and compliance are not provided as a single managed service by the model itself.
Deployment and API availability
OLMoASR-small.en is intended for local or self-managed inference. Ai2 documents Python-based use through the olmoasr package, where a user loads the model and invokes its transcription functionality on an audio file.
No official first-party hosted inference price is specified for this exact checkpoint. There is also no documented managed API, batch API, streaming interface, or hosted per-minute billing model in the supplied information. Consequently, the practical cost is determined by the infrastructure used to run it, including compute, storage, operations, and engineering time.
The 244-million-parameter size gives the model a larger deployment footprint than the 39-million-parameter OLMoASR-tiny.en and 74-million-parameter OLMoASR-base.en variants. It will generally require more memory and compute than those smaller checkpoints. In exchange, Ai2 reports lower word error rates than the smaller variants. The supplied documentation does not specify a context window, maximum output-token limit, minimum hardware configuration, or a fixed real-time factor, so those values should not be assumed.
Capabilities that are not part of the model
OLMoASR-small.en is an ASR checkpoint rather than a general-purpose assistant or language model. Its role is to convert audio into text, not to reason over a conversation, write software, browse the web, call external tools, or produce structured application actions.
- Reasoning: No general reasoning capability is documented. Any analysis of the transcript must be performed by a separate model or application.
- Coding: No code-generation capability is documented.
- Tool and function calling: Not supported or documented.
- Web search: Not supported.
- Structured output: No structured-output mode is documented, although the transcription includes sentence-level timestamp information.
- Streaming: No streaming interface is documented for this exact checkpoint.
A transcription workflow can still be combined with other software. For example, an application could pass OLMoASR-small.en's transcript to a separate language model for summarization or extraction. That would be a multi-component system, however, and should not be attributed to OLMoASR-small.en itself.
Strengths and limitations
Main strengths
- Open deployment: The weights and supporting project materials are available for self-managed use under the stated Apache 2.0 release.
- Useful middle position: The small checkpoint provides a compromise between the lower resource requirements of tiny and base models and the higher resource requirements of medium and large models.
- Long-form support: It is designed for extended recordings as well as short clips.
- Timestamped results: Sentence-level timestamps support captions, search, indexing, and media navigation.
- Research transparency: Ai2 releases training, processing, evaluation, and model-related artifacts rather than offering only a closed transcription endpoint.
Main limitations
- English only: It is not a suitable primary choice for multilingual transcription.
- Self-managed operations: Users must arrange hosting and inference instead of calling a documented first-party paid API.
- Resource requirements: The model is larger and more demanding than the tiny and base checkpoints.
- Variable real-world accuracy: Noise, overlapping speakers, accents, poor microphones, and unfamiliar domains can reduce transcription quality.
- Narrow task scope: It does not synthesize speech, generate images or video, reason over content, or provide application tools.
- Privacy responsibilities: Organizations must address consent, copyright, confidentiality, and data-protection obligations when processing recordings.
When to choose OLMoASR-small.en
Choose OLMoASR-small.en when you need an open English transcription model and can operate it locally or on infrastructure you control. It is a reasonable fit for meeting archives, podcast and lecture transcription, timestamped media search, captioning pipelines, call-recording analysis, and ASR experimentation where inspectable model artifacts matter.
Its middle position in the OLMoASR family is particularly useful when the smallest checkpoints do not provide enough accuracy, but the compute requirements of the medium or large variants are difficult to justify. The reported results also make it a relevant candidate for users comparing open speech recognition with similarly sized proprietary or open alternatives.
Another option may be more appropriate in several situations. Select a smaller OLMoASR variant when lower memory use, faster inference, or a smaller deployment footprint is more important than the reported accuracy trade-off. Consider a larger OLMoASR variant when accuracy is the priority and additional compute is available. Choose a managed transcription service when you need a hosted endpoint, predictable operational support, or usage-based billing rather than self-managed inference. Use a multilingual ASR system when the recordings contain languages beyond English.
Pricing and practical cost
There is no official hosted API price supplied for OLMoASR-small.en. The open checkpoint can be downloaded for self-managed use, but “open” does not mean cost-free in production: users may still pay for GPUs or other compute, storage, networking, monitoring, and engineering support.
Because no fixed hosted price, context limit, maximum output limit, or streaming guarantee is documented for this exact model, teams should benchmark it on representative audio before estimating throughput or total cost. A useful evaluation should include the recording lengths, accents, noise levels, speaker overlap, and vocabulary found in the intended application.
Bottom line
OLMoASR-small.en is a focused, open English speech-recognition model for users who want local control, timestamped transcription, and a balance between model size and reported accuracy. Its value is strongest in self-managed transcription and research workflows. It is not a hosted conversational assistant or a general-purpose multimodal model, and its English-only scope, infrastructure requirements, and lack of a documented first-party API should be considered before deployment.

