OLMoASR

OLMoASR-medium.en

An open-weight Ai2 speech recognition model for English short- and long-form transcription, with sentence-level timestamps and no official hosted API price.

Text Reasoning Coding
OLMoASR-medium.en is an open-weight speech recognition model from the Allen Institute for Artificial Intelligence (Ai2). Released on August 28, 2025, it is designed to convert English speech into text for meetings, calls, lectures, podcasts, broadcasts, audiobooks, and other real-world recordings. Its main practical distinction is that users can download and run the checkpoint themselves, while Ai2 does not publish an official hosted per-minute or per-token price for this exact model.
Outputs

What OLMoASR-medium.en can produce

Text
Inputs

What it can understand

Audio
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
6/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family OLMoASR
Model type Other
Release date 2025-08-28
Status Current open-weight model
Knowledge cutoff notes

Knowledge-cutoff metadata is not applicable or publicly documented for this speech recognition checkpoint. The model is trained for audio transcription rather than open-ended factual text generation.

Model notes

OLMoASR-medium.en is the 769-million-parameter medium variant in Ai2's OLMoASR family. It is trained on English-only audio-text data and uses a Transformer encoder-decoder design. The official implementation supports Python inference through the olmoasr package and returns transcription text, segments, timestamps, tokens, and language metadata. Documentation describes 30-second audio chunks and sentence-level timestamps. Ai2 reports average WER of 12.8% on its short-form evaluation suite and 11.0% on its long-form suite. The checkpoint is open-weight under Apache 2.0. No official hosted per-minute or per-token price is documented for this exact model.

Cost

Model pricing

Input No official hosted API price; self-hosted checkpoint
Output No official hosted API price; self-hosted checkpoint
Model guide

OLMoASR-medium.en: Ai2’s Open English Speech-to-Text Model

OLMoASR-medium.en is Ai2’s 769-million-parameter, English-only automatic speech recognition model for short- and long-form transcription. It produces text with sentence-level timestamps and is distributed as an open-weight Apache 2.0 checkpoint for local deployment rather than as a priced hosted API.

What is OLMoASR-medium.en?

OLMoASR-medium.en is an English automatic speech recognition (ASR) model developed by Ai2 as part of the OLMoASR family. ASR systems turn spoken audio into written text; unlike a general-purpose language model, this model is primarily designed to transcribe recordings rather than hold open-ended conversations or generate arbitrary text.

The medium variant contains approximately 769 million parameters. It uses a Transformer-based encoder-decoder design: an audio encoder processes the speech signal and a language decoder produces the transcription. The published implementation is intended for local or user-managed inference through the olmoasr Python package.

OLMoASR-medium.en fits into Ai2’s broader open-model research catalog, but it has a narrower role than Ai2’s general language and multimodal models. Its output is text derived from audio, with additional transcription metadata such as segments and timing information.

Capabilities and input/output types

The model accepts audio and returns text. It does not provide direct audio, image, or video output. The documented implementation can return a complete transcription as well as individual segments, start and end timestamps, token information, language identification, and decoding statistics.

OLMoASR-medium.en supports both short-form and long-form English speech recognition. The published feature-extraction configuration uses a 16 kHz sampling rate and 30-second audio chunks. Chunking is an implementation detail rather than a statement that the complete recording must be limited to 30 seconds; the model is specifically described as supporting long-form transcription.

Useful applications include transcribing meetings and telephone calls, generating lecture captions, processing podcasts and audiobooks, indexing broadcasts, and extracting text for speech analytics. Sentence-level timestamps make the output more useful for captioning, review interfaces, and locating spoken passages within a recording.

Reported performance

Ai2 reports an average word error rate (WER) of 12.8% across its short-form evaluation suite and 11.0% across its long-form evaluation suite. Word error rate measures transcription mistakes, with lower values generally indicating more accurate recognition. The reported evaluations covered domains including audiobooks, telephone conversations, meetings, lectures, and accented speech.

Ai2 also reports that OLMoASR-medium.en performed in the same general range as Whisper-medium.en at a similar parameter count on the cited benchmark comparison. This is a provider-reported evaluation claim, not a guarantee for every language variety, recording environment, microphone, speaker, or background-noise condition. Actual results can vary substantially with audio quality and the characteristics of the material being transcribed.

Availability, licensing, and pricing

The model checkpoint is available through Ai2’s OLMoASR resources on Hugging Face, and the source repository and documentation describe Python-based inference. The documented loading interface uses the olmoasr package and load_model("medium", inference=True). Command-line support was identified as still being in development in the documented usage instructions.

OLMoASR-medium.en is licensed under Apache 2.0 according to the supplied model information. This makes it possible to deploy the checkpoint in a user-controlled environment subject to the license and any applicable usage requirements.

There is no official hosted inference price listed by Ai2 for this exact checkpoint. It should therefore not be described as free hosted transcription: running it generally requires the user to supply compatible compute, storage, audio-processing infrastructure, and operational support. A third-party host may charge for those resources, but such charges are separate from the model’s published checkpoint licensing and cannot be inferred from Ai2’s materials.

Limits and undocumented features

The most important functional limitation is language coverage: OLMoASR-medium.en is English-only. It is not a documented multilingual transcription model, so users needing native support for multiple languages should consider an option specifically evaluated and released for that purpose.

The supplied documentation does not specify a token-based context window or maximum generated-token limit. Those concepts are less directly useful for a speech transcription checkpoint than for a conversational text model, but the absence of a published value means no precise maximum should be assumed. The documented processing setup does identify 30-second audio chunks, while the model’s stated use includes long-form transcription.

The supplied research also does not document a structured-output interface, prompt caching, batch API, fine-tuning support, hosted web-search access, or a first-party streaming API. It does not provide tool or function calling. It should be treated as a transcription model rather than an agent platform or general-purpose developer API.

Reasoning, coding, and tool support

OLMoASR-medium.en does not perform reasoning in the usual language-model sense. It decodes speech into text and may provide transcription metadata, but it is not intended to solve multi-step problems, answer questions about a recording, or make independent decisions.

It also has no documented coding capability beyond transcribing spoken code or technical discussion as audio. It cannot be relied on to generate, execute, test, or explain software. Similarly, there is no documented tool-use or function-calling mechanism. Any workflow that needs summarization, question answering, translation, redaction, search, or structured database actions would need additional software or another model after transcription.

Speed and cost trade-offs

OLMoASR-medium.en’s cost profile is different from that of a hosted speech API. There is no published per-minute charge for the exact model, so per-recording cost depends on the hardware and inference setup chosen by the operator. Self-hosting can be attractive for high-volume or privacy-sensitive workloads when the organization already has suitable infrastructure, but it transfers responsibility for deployment, scaling, monitoring, and performance tuning to the user.

Ai2 does not publish a universal throughput or latency figure in the supplied material. Processing speed therefore cannot be stated independently of hardware, audio duration, batch configuration, and implementation details. The model’s 769-million-parameter size represents a meaningful local-compute requirement compared with very small transcription models, but it should not be converted into a specific memory, latency, or cost estimate without a documented test environment.

In practical terms, a managed transcription service may be simpler for occasional workloads or teams that need predictable operational controls. OLMoASR-medium.en may be more suitable when open weights, local execution, customization, and control over the deployment environment matter more than turnkey hosting.

Main strengths and limitations

Strengths

  • Open deployment model: the checkpoint is available for user-managed inference rather than being limited to a closed hosted endpoint.
  • Long- and short-form support: it is designed for both brief clips and longer recordings.
  • Useful timing metadata: sentence-level timestamps and segment information support captions, review, and searchable archives.
  • Broad English evaluation coverage: Ai2 reports testing across calls, meetings, lectures, audiobooks, accented speech, and other domains.
  • Apache 2.0 licensing: the stated license is familiar to many research and commercial engineering teams, subject to reviewing the applicable terms.

Limitations

  • English only: it is not positioned as a multilingual ASR solution.
  • No official hosted price: users must arrange their own compute or use a separately configured third-party service.
  • Limited product layer: the model itself does not provide chat, summarization, search, tool calls, or workflow automation.
  • Undocumented operational limits: official material supplied here does not specify throughput, latency, context length, maximum output tokens, or a universal hardware requirement.
  • Deployment responsibility: self-hosting requires engineering work for installation, scaling, monitoring, and integration.

When to choose OLMoASR-medium.en

Choose OLMoASR-medium.en when the primary requirement is English speech-to-text and you value an open checkpoint that can be run in your own environment. It is a reasonable candidate for research, local transcription pipelines, meeting and lecture archives, podcast processing, broadcast indexing, and applications that need sentence-level timing without depending on a proprietary transcription endpoint.

It is especially worth considering when the team can manage its own infrastructure and wants more control over data handling, model artifacts, and deployment than a hosted API typically provides. The reported short-form and long-form WER figures also make it relevant for users comparing open ASR models on varied English audio, although those results should be validated on the intended recordings.

Another option may be more appropriate when multilingual recognition is essential, when a guaranteed hosted price and service-level operation are required, or when the workflow needs built-in diarization, translation, summarization, search, streaming, or downstream actions that are not documented capabilities of this checkpoint. A general language model may be useful after transcription for analysis, but it is not a substitute for the audio recognition stage itself.

Bottom line

OLMoASR-medium.en is a focused, open-weight English transcription model rather than a complete speech platform or conversational assistant. Its useful combination of long-form support, sentence-level timestamps, reported evaluation across varied domains, and local deployment makes it a practical option for teams that want control over the ASR layer. The trade-off is that pricing, speed, hosting, scaling, multilingual coverage, and all post-transcription processing remain outside the checkpoint and must be handled separately.


Answers to Frequently Asked Questions

What are the main limitations of OLMoASR-medium.en?
OLMoASR-medium.en is English-only and is designed for transcription rather than conversation, reasoning, coding, summarization, translation, search, tool use, or workflow automation. It also has no documented first-party streaming API, hosted service price, universal hardware requirement, or guaranteed throughput and latency figure.
Is OLMoASR-medium.en free to use, and what license does it have?
The model checkpoint is available under the Apache 2.0 license for user-managed inference, subject to the applicable terms. Ai2 does not list an official hosted inference price for this checkpoint, so users must provide compatible compute and infrastructure or use a separately priced third-party host.
How accurate is OLMoASR-medium.en?
Ai2 reports an average word error rate of 12.8% on its short-form evaluation suite and 11.0% on its long-form evaluation suite. Actual accuracy can vary depending on the recording quality, accents, background noise, speakers, and audio domain.
What is OLMoASR-medium.en used for?
OLMoASR-medium.en is an English automatic speech recognition model for converting spoken audio into written text. It can be used for meetings, telephone calls, lectures, podcasts, audiobooks, broadcasts, captions, and speech analytics.
Does OLMoASR-medium.en support long audio recordings?
Yes. OLMoASR-medium.en supports both short-form and long-form English transcription. Its implementation uses 16 kHz audio processing and 30-second chunks, but the model is not limited to recordings of only 30 seconds.


Sources 5
Provider

About Allen Institute for Artificial Intelligence (Ai2)