OLMoASR

OLMoASR-large.en-v1

by Allen Institute for Artificial Intelligence (Ai2) · Available open-weight model

An open 1.5-billion-parameter English speech recognition model from Ai2, trained on 440,000 hours of audio and released with weights, code, data-processing tools, and evaluation resources.

Text Reasoning Coding
OLMoASR-large.en-v1 is the first large-scale model in Ai2’s OLMoASR family of open English speech recognition systems. Released on August 28, 2025, it was trained from scratch on 440,000 hours of audio and is designed for zero-shot transcription across varied speech domains, including conversations, meetings, lectures, audiobooks, and calls.
Outputs

What OLMoASR-large.en-v1 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family OLMoASR
Model type Other
Release date 2025-08-28
Status Available open-weight model
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this speech recognition checkpoint. The model is trained for audio transcription rather than general factual question answering.

Model notes

OLMoASR-large.en-v1 is the 1.5-billion-parameter large v1 checkpoint in Ai2's OLMoASR family. The official release identifies the checkpoint as OLMoASR-large.en-v1, while the Hugging Face repository stores the weights under the OLMoASR-large.en directory. It was trained on 440,000 hours of audio. Ai2 reports average WER of 13.0% on short-form benchmarks and 11.4% on long-form benchmarks. The model is English-focused, produces transcription text and timestamped segments, and is intended for local Python inference. Editorial scores are comparative estimates for a speech recognition model and are not vendor ratings.

Cost

Model pricing

Input No official hosted API pricing; released for open-weight deployment
Output No official hosted API pricing; released for open-weight deployment
Model guide

OLMoASR-large.en-v1: Ai2’s Open 1.5B English Speech Recognition Model

OLMoASR-large.en-v1 is a fully open, 1.5-billion-parameter English automatic speech recognition model from the Allen Institute for AI. It supports short- and long-form transcription, sentence-level timestamps, and local deployment with openly released weights, code, training data, and evaluation tools.

What is OLMoASR-large.en-v1?

OLMoASR-large.en-v1 is an open automatic speech recognition model developed by the Allen Institute for AI, also known as Ai2. Automatic speech recognition, or ASR, converts spoken audio into written text. This checkpoint is focused on English transcription and contains approximately 1.5 billion parameters.

The model is intended primarily for local or research deployment rather than use through a metered, provider-managed API. Ai2 released the model with openly available weights and supporting research resources, including code, data-processing tools, and evaluation code. The released checkpoint is commonly represented in the model repository by the directory OLMoASR-large.en; the -v1 designation distinguishes it from the later OLMoASR-large.en-v2 checkpoint.

Within Ai2’s current open-model catalog, OLMoASR-large.en-v1 is the speech recognition specialist in the OLMo family. It is separate from Ai2’s text-focused OLMo language models and Molmo multimodal models. Its job is narrower but clearer: turning English speech into searchable, editable, and time-aligned text.

Core capabilities and output

OLMoASR-large.en-v1 accepts English audio and produces transcription text. It supports both short-form utterances and long-form recordings, making it suitable for anything from a brief voice clip to a meeting, lecture, call, audiobook, or podcast-style recording.

The implementation can return sentence-level timestamp information alongside the recognized text. This allows an application to associate each sentence with a start and end position in the source audio. The documented transcription result also includes segment boundaries, token information, confidence-related generation statistics, and detected-language information.

  • English speech-to-text transcription
  • Short-form and long-form audio processing
  • Sentence-level timestamps
  • Local Python inference
  • Streaming-oriented use cases documented by Ai2
  • Open model weights, training pipeline, data-processing code, and evaluation code

The model produces text and transcription metadata. It does not generate speech, music, images, video, embeddings, or executable actions. It should therefore be evaluated as an ASR component rather than as a general-purpose conversational assistant.

Training and release context

Ai2 announced the OLMoASR family on August 28, 2025. OLMoASR-large.en-v1 was trained on 440,000 hours of audio. Ai2 describes the family as being trained from scratch using curated public audio and transcript data rather than undisclosed proprietary training sources.

The broader OLMoASR project includes a weakly supervised pool of approximately three million hours of English audio and a filtered OLMoASR-Mix dataset of approximately one million hours. Ai2 also released the data-processing and filtering pipeline. For researchers, this openness is important because it makes it easier to inspect how data curation affects model behavior and to reproduce or extend the training process.

The later large.en-v2 checkpoint was trained on 680,000 hours of audio. That comparison helps position v1 within the family, but it does not make v1 interchangeable with v2: OLMoASR-large.en-v1 is a distinct released checkpoint with its own training data and benchmark results.

Recognition performance

Ai2 reported an average word error rate, or WER, of 13.0 percent for OLMoASR-large.en-v1 on its short-form evaluation suite and 11.4 percent on its long-form evaluation suite. WER measures the number of word-level insertions, deletions, and substitutions relative to a reference transcript. Lower is better, but an average benchmark score is not a guaranteed error rate for a particular recording.

Results varied according to the dataset and speech conditions. Clean speech was easier for the model than highly noisy, distant, overlapping, or far-field conversational recordings. Accents, microphone quality, background noise, speaker overlap, and transcript conventions can all affect the practical result.

In Ai2’s cited comparison, OLMoASR-large.en-v1 achieved a 13.0 percent short-form average WER versus 12.2 percent for Whisper-large-v1. The later OLMoASR-large.en-v2 narrowed that gap after training on more audio. These figures are useful for positioning the model, but they should not be treated as a universal ranking across every language, domain, audio condition, or deployment configuration.

Deployment, pricing, and resource trade-offs

There is no official hosted API price identified for OLMoASR-large.en-v1. Ai2 released it for open-weight deployment, so the direct model price is not a recurring per-minute or per-token subscription. The practical cost comes from the infrastructure needed to run inference, along with engineering, storage, and maintenance costs.

The official materials document Python-based use through the OLMoASR codebase and require its dependencies and audio-processing tools such as FFmpeg. The model can be loaded from Ai2’s model collection and used to transcribe an audio file. Exact hardware requirements depend on precision, batch size, recording duration, and the chosen inference setup.

As the large checkpoint in the v1 family, it is expected to require more resources than the tiny, base, small, and medium OLMoASR variants. The supplied specifications do not publish a fixed context length or maximum output-token limit. For an ASR system, practical limits are more closely tied to audio duration, memory, batching, and long-form processing behavior than to a conventional language-model context-window figure.

SpecificationVerified information
ProviderAllen Institute for AI (Ai2)
Release dateAugust 28, 2025
Model sizeApproximately 1.5 billion parameters
Primary inputEnglish audio
Primary outputTranscription text and timestamped segments
Training audio440,000 hours
Hosted API priceNo official hosted API pricing identified
Context lengthNot published in the supplied specifications
Maximum output tokensNot published in the supplied specifications

Main strengths and limitations

Strengths

  • Open research and deployment: The weights, code, data-processing pipeline, and evaluation resources are available, which is useful for organizations that need local control or want to study the system.
  • Useful long-form support: The model is designed for more than isolated voice commands and can process recordings such as meetings, lectures, calls, and audiobooks.
  • Timestamped results: Sentence-level timing helps with captions, searchable archives, editing workflows, and review interfaces.
  • English specialization: Focusing on English allows the checkpoint to target English transcription rather than attempting to cover many languages with one model.
  • No required per-request provider fee: Local deployment can avoid hosted inference charges when an organization already has suitable infrastructure.

Limitations

  • English-focused: It is not a general multilingual transcription model. A multilingual ASR system is more appropriate when the application must support several languages.
  • Self-hosting responsibility: Users must manage hardware, dependencies, audio preprocessing, scaling, monitoring, and updates. An open checkpoint is not the same as a turnkey hosted service.
  • Variable accuracy: Noisy, distant, overlapping, heavily accented, or poorly aligned speech can produce substantially worse transcripts than clean recordings.
  • No general assistant features: The model does not provide built-in web search, tool calling, structured-output controls, or a provider-managed conversational interface.
  • No speech generation: It recognizes speech but does not synthesize a spoken response.
  • Unpublished fixed limits: The supplied documentation does not specify a conventional context window, maximum output-token count, or a hosted service quota.

Reasoning, coding, and tool support

OLMoASR-large.en-v1 is not a reasoning or coding model in the usual language-model sense. Its output is transcription text and related metadata, not step-by-step analysis, source code, or autonomous decisions. Any reasoning or coding capability in an application would need to come from a separate model used after transcription.

Similarly, the model has no documented tool or function-calling interface. It can serve as one stage in a larger workflow—for example, transcribing a meeting before another system summarizes it—but the ASR checkpoint itself does not browse the web, call external functions, execute code, or perform actions.

When to choose OLMoASR-large.en-v1

Choose OLMoASR-large.en-v1 when the priority is open, local English transcription and the team can operate its own inference environment. It is a reasonable candidate for research on speech recognition, internal meeting transcription, lecture and call archives, captioning pipelines, podcast processing, and speech-data experiments where inspectable model artifacts matter.

Its openness is especially valuable when audio should remain inside an organization’s infrastructure or when researchers need access to the training and evaluation pipeline. It can also be attractive when avoiding a provider’s per-minute API pricing is more important than minimizing operational work.

Another option may be more appropriate in several situations. Use a multilingual ASR model for multilingual workloads, a hosted transcription API when infrastructure simplicity and managed scaling are more important, or a smaller OLMoASR variant when lower hardware use and faster inference matter more than the capabilities of the large checkpoint. A later OLMoASR-large.en-v2 checkpoint may be worth evaluating when its larger training corpus and reported results better match the target workload.

For highly noisy or overlapping speech, benchmark the model on representative recordings before deployment. The reported WER averages are useful reference points, but production quality should be measured on the accents, microphones, environments, and vocabulary that the application will actually encounter.

Bottom line

OLMoASR-large.en-v1 is a research-oriented, open-weight English ASR model rather than a general AI assistant or hosted transcription product. Its defining advantages are local deployment, transparent supporting resources, long-form transcription, and timestamped output. Its main trade-offs are English-only focus, self-hosting complexity, resource requirements, and the absence of a managed API with published usage limits and prices.

For teams that value reproducibility and control, it offers a substantial open checkpoint for English speech-to-text work. For teams that need multilingual coverage, turnkey scaling, speech generation, or integrated assistant features, a different model type or service will be a better fit.


Answers to Frequently Asked Questions

Who should use OLMoASR-large.en-v1?
It is best suited to teams and researchers seeking open, local English transcription with inspectable weights, code, data-processing tools, and evaluation resources. A multilingual ASR model, hosted transcription API, smaller model, or later OLMoASR-large.en-v2 checkpoint may be more appropriate for other requirements.
Does OLMoASR-large.en-v1 have an official hosted API price?
No official hosted API pricing has been identified for OLMoASR-large.en-v1. It is released for open-weight, local deployment, so users avoid a required per-minute or per-token model fee but must cover infrastructure, storage, engineering, and maintenance costs.
How accurate is OLMoASR-large.en-v1?
Ai2 reported an average word error rate (WER) of 13.0% on its short-form evaluation suite and 11.4% on its long-form evaluation suite. Actual accuracy can vary significantly with accents, background noise, microphone quality, distant or overlapping speech, and transcript conventions.
What is OLMoASR-large.en-v1?
OLMoASR-large.en-v1 is an open-weight automatic speech recognition model developed by the Allen Institute for AI (Ai2). It contains approximately 1.5 billion parameters and converts English speech into written text.
What types of audio and output does OLMoASR-large.en-v1 support?
The model supports short-form and long-form English audio, including meetings, lectures, calls, audiobooks, and podcasts. It produces transcription text and metadata such as sentence-level timestamps, segment boundaries, token information, and confidence-related generation statistics.


Sources 4
Provider

About Allen Institute for Artificial Intelligence (Ai2)