OLMoASR

OLMoASR-large.en-v2

by Allen Institute for Artificial Intelligence (Ai2) · Current open-weight model; downloadable and usable for local inference

Ai2’s OLMoASR-large.en-v2 is an open-weight 1.5-billion-parameter English speech-recognition model trained on approximately 680,000 hours of audio. It supports short- and long-form transcription, sentence-level timestamps, local Python inference, and reproducible research workflows, but has no official Ai2 hosted API pricing and requires more deployment resources than smaller ASR models.

Text Reasoning Coding
OLMoASR-large.en-v2 is the largest model in Ai2’s initial OLMoASR speech-recognition family. It is designed for English audio transcription rather than conversation, speech generation, or general-purpose text production. The model accepts audio, returns text with segment timing information, and can be run locally through the project’s Python-based implementation. Its main distinction is openness: Ai2 provides not only model weights but also training and evaluation resources that help researchers inspect, reproduce, and adapt the system.
Outputs

What OLMoASR-large.en-v2 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
5/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family OLMoASR
Model type Other
Release date 2025-08-28
Status Current open-weight model; downloadable and usable for local inference
Knowledge cutoff notes

A language-model knowledge cutoff is not applicable to this specialized automatic speech recognition checkpoint. Ai2 does not publish a conventional knowledge-cutoff date for the exact model.

Model notes

OLMoASR-large.en-v2 is a specialized English ASR checkpoint in the OLMoASR family. It has approximately 1.5 billion parameters and was trained on approximately 680,000 hours of audio. The project supports short-form and long-form transcription and returns sentence-level timestamps. The reference implementation documents Python inference and currently does not provide an official Ai2 hosted per-minute API price. The model is distributed through the Ai2 OLMoASR repository and Hugging Face model collection. Editorial scores are comparative estimates for specialized ASR use, not provider-published ratings.

Cost

Model pricing

Input No official hosted API pricing; self-hosted open weights
Output No official hosted API pricing; self-hosted open weights
Model guide

OLMoASR-large.en-v2: Ai2’s Open 1.5B Model for English Speech Transcription

OLMoASR-large.en-v2 is Ai2’s open-weight English automatic speech recognition model. With approximately 1.5 billion parameters, it transcribes short and long recordings, produces sentence-level timestamps, and was trained on approximately 680,000 hours of audio. Its open weights, code, data pipelines, and evaluation tools make it a useful option for self-hosted transcription and reproducible speech-recognition research.

What OLMoASR-large.en-v2 is

OLMoASR-large.en-v2 is an open-weight automatic speech recognition (ASR) model from Ai2, the Allen Institute for AI. ASR systems convert spoken language in an audio recording into written text. This checkpoint is focused on English and is intended for both short utterances and long-form recordings such as meetings, lectures, interviews, calls, and podcasts.

The model contains approximately 1.5 billion parameters and is the large v2 checkpoint in Ai2’s initial OLMoASR family. Ai2 released it on August 28, 2025, alongside other OLMoASR variants and supporting project materials. Unlike a hosted transcription service, it is primarily a model artifact that users download and run themselves or access through a compatible third-party deployment.

Its output is transcription text with sentence-level timestamps. These timestamps identify where recognized segments begin and end in the recording, which is useful for captions, searchable archives, transcript navigation, and reviewing uncertain sections of an audio file.

Capabilities and supported modalities

OLMoASR-large.en-v2 has a narrow but clearly defined role: English speech-to-text transcription. Its supported input is audio, and its direct model output is text accompanied by segment metadata such as start and end times.

  • English automatic speech recognition
  • Short-form and long-form audio transcription
  • Sentence-level timestamps
  • Local Python-based inference
  • Open-weight deployment and research workflows

The model does not generate audio, images, video, embeddings, or general-purpose conversational responses. It is also not an image-and-audio multimodal assistant: the supplied specifications identify audio as its input modality, with text as its output. There is no documented tool or function-calling interface, web search capability, JSON mode, or structured-output feature.

Because it is a specialized ASR checkpoint, conventional language-model fields such as context-window length and maximum output-token count are not published for this model. Transcription capacity should instead be understood in terms of the audio files and processing workflow supported by the reference implementation.

Accuracy and benchmark results

Ai2 evaluated the OLMoASR family on 21 speech-recognition benchmarks covering categories such as audiobooks, telephone calls, meetings, lectures, and conversational speech. For OLMoASR-large.en-v2, the reported average word error rate was 12.6% across the short-form benchmarks and 11.5% across the long-form benchmarks.

Word error rate, or WER, measures the proportion of recognized words that differ from the reference transcript. Lower is better, but an average does not predict performance on every recording. The reported short-form results illustrate the model’s variation by domain:

BenchmarkReported word error rate
LibriSpeech clean2.7%
LibriSpeech other5.6%
TED-LIUM34.2%
WSJ3.6%
CallHome15.0%
Switchboard11.7%
Common Voice 5.111.1%

These figures are provider-reported evaluation results, not a guarantee for an individual recording. Recognition can vary with background noise, microphone quality, speaker accents, overlapping speech, vocabulary, and the subject area being discussed. Telephone conversations and informal speech may be more difficult than clean, prepared recordings.

Training and what is open

Ai2 says that OLMoASR-large.en-v2 was trained from scratch using the OLMoASR-Mix dataset. That dataset was distilled from a larger weakly supervised audio-text pool of approximately three million hours. The curation process included language alignment, transcript-quality filtering, deduplication, and other heuristics intended to improve the reliability of audio-transcript pairs.

The project’s openness extends beyond the checkpoint itself. Public materials include the model implementation, training procedures, data-processing code, evaluation scripts, and associated dataset resources. This makes the model relevant to researchers who need to inspect the development pipeline rather than rely exclusively on a proprietary hosted endpoint.

For developers, this can support experiments with reproducible evaluation, domain-specific testing, auditing, and self-hosted transcription. It also means that using the model is more involved than sending an audio file to a commercial API: users must obtain the code and weights, install dependencies, prepare suitable hardware, and manage the inference environment.

Deployment, pricing, and performance trade-offs

The reference project documents Python inference and requires the OLMoASR codebase, its dependencies, and audio-processing software such as ffmpeg. The approximately 1.5-billion-parameter size makes this checkpoint substantially more demanding to deploy than smaller OLMoASR variants. The supplied research does not specify a minimum GPU, CPU, memory requirement, throughput figure, or latency target, so those details should be tested for the intended hardware rather than assumed.

There is no official Ai2 hosted per-minute API price for OLMoASR-large.en-v2. Its direct cost model is therefore different from a usage-priced transcription API: the weights are available for download, while the operator pays for hardware, storage, electricity, maintenance, and engineering time when running it locally. A third-party host may charge for inference, but that would be a separate provider’s pricing rather than an Ai2 price for the model.

Its speed and cost trade-off are also different from those of a hosted service. Local execution can offer control over data handling and predictable access after deployment, but the large checkpoint may require more compute and may be slower or more expensive to operate than a smaller speech-recognition model. The research supplies no official speed ranking or cost benchmark. Any comparative speed or cost assessment should therefore be treated as deployment-specific.

Main strengths and limitations

Strengths

  • Open model artifacts: Users can access weights and project resources rather than depending only on a closed endpoint.
  • Reproducibility: Training, data-processing, and evaluation materials support more transparent research and auditing.
  • Long-form support: The model is designed for extended recordings as well as short utterances.
  • Timing information: Sentence-level timestamps make the output more useful for captions, archives, and transcript navigation.
  • Broad evaluation coverage: Ai2 reports testing across 21 benchmarks and several speech domains.

Limitations

  • English focus: It should not be selected as a multilingual transcription model.
  • Specialized output: It produces transcriptions, not speech, images, video, embeddings, or general conversational answers.
  • Hardware burden: The large parameter count makes it less suitable for highly resource-constrained deployments than smaller family members.
  • No official Ai2 hosted API: Users generally need to self-host the model or find a compatible external deployment.
  • Variable real-world accuracy: Noisy, overlapping, accented, poorly recorded, or specialized speech may reduce transcription quality.
  • No documented built-in tools: The model does not provide web search, function calling, or an assistant-style application layer.

When to choose OLMoASR-large.en-v2

Choose OLMoASR-large.en-v2 when the primary requirement is open, self-hosted English transcription and the team can support a relatively large local model. It is a strong fit for research groups benchmarking ASR systems, organizations building searchable audio archives, developers creating meeting or lecture transcription workflows, and projects that need more control over where audio is processed.

The model is also a reasonable choice when transparency matters. Access to the implementation, training pipeline, data-processing resources, and evaluation code can be more valuable than a simpler commercial API for researchers studying model behavior or building reproducible experiments.

Another option may be more appropriate when the priority is very low latency, minimal hardware, multilingual support, or turnkey hosted operation. A smaller OLMoASR checkpoint may be preferable for edge devices and resource-constrained servers. A commercial transcription API may be easier for applications that need managed scaling, published per-minute billing, service-level support, or an integration that does not require maintaining model infrastructure. Those alternatives may sacrifice some of the inspectability and deployment control that distinguish Ai2’s model.

Reasoning, coding, and application fit

OLMoASR-large.en-v2 should not be evaluated like a general language model. It does not perform open-ended reasoning or code generation as a primary capability. It can be incorporated into a software workflow through local Python inference, but that is an integration method rather than evidence that the model writes or executes code. Similarly, timestamps and transcript metadata do not constitute general structured-output or agent capabilities.

Editorial capability ratings supplied for the model assign little practical meaning to reasoning and coding because those tasks are outside its purpose. They should not be read as provider-published benchmark scores. The useful evaluation questions are instead whether its English transcripts are accurate enough for the target audio, whether its timestamps meet the application’s needs, and whether the available hardware can process recordings at an acceptable rate.

Bottom line

OLMoASR-large.en-v2 is a focused English ASR checkpoint for users who value open weights, transparent research materials, and local deployment. Its approximately 1.5-billion-parameter design, long-form transcription support, sentence-level timestamps, and reported benchmark results make it a serious research and self-hosting option. Its limitations are equally important: it is not multilingual, it is not a conversational model, it has no official Ai2 hosted API pricing, and its large size increases operational demands. The best reason to choose it is control and reproducibility, not access to a broad assistant feature set.


Answers to Frequently Asked Questions

Who should choose OLMoASR-large.en-v2?
It is best suited to researchers, developers, and organizations that need open, self-hosted English transcription, sentence-level timestamps, and transparent training and evaluation materials. A smaller model or commercial transcription API may be more suitable when low hardware requirements, multilingual support, managed scaling, or turnkey operation is more important.
Can OLMoASR-large.en-v2 be used locally, and does Ai2 offer an official hosted API?
Yes. The model is intended for local Python-based inference using the OLMoASR codebase, its dependencies, and audio-processing tools such as ffmpeg. Ai2 does not provide an official hosted per-minute API price for this model, so users generally self-host it or use a separate third-party deployment.
How accurate is OLMoASR-large.en-v2?
Ai2 reports an average word error rate of 12.6% on short-form benchmarks and 11.5% on long-form benchmarks. Results vary by recording conditions and domain: reported word error rates range from 2.7% on LibriSpeech clean to 15.0% on CallHome.
What is OLMoASR-large.en-v2?
OLMoASR-large.en-v2 is an open-weight English automatic speech recognition model from Ai2, the Allen Institute for AI. It contains approximately 1.5 billion parameters and converts audio into text with sentence-level timestamps.
What types of audio can OLMoASR-large.en-v2 transcribe?
The model supports both short utterances and long-form English recordings, including meetings, lectures, interviews, calls, podcasts, and audiobooks. Its output includes transcript text and the start and end times of recognized sentence-level segments.


Sources 5
Provider

About Allen Institute for Artificial Intelligence (Ai2)