Granite Speech 5.0

Granite Speech 5.0 470M TurboCTC NC

by IBM watsonx · Current; research and noncommercial use only

A compact IBM English speech recognition checkpoint built with a Conformer encoder and CTC decoding. It is optimized for fast local transcription and edge-oriented inference, but its CC-BY-NC-SA-4.0 license restricts the model to research and noncommercial use.

Text Reasoning Coding
Granite Speech 5.0 470M TurboCTC NC is IBM's noncommercial fast-transcription model for English audio. It converts speech into text with a non-autoregressive CTC decoder, making it suitable for local inference, batch processing, browser demonstrations, and latency-sensitive research projects. The main trade-off is licensing: this checkpoint is intended for research and noncommercial use, not commercial production deployment.
Outputs

What Granite Speech 5.0 470M TurboCTC NC can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
10/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite Speech 5.0
Model type Other
Release date 2026-08-25
Status Current; research and noncommercial use only
Knowledge cutoff notes

The model documentation does not specify a knowledge cutoff. As an encoder-only speech recognition model, it is not primarily a knowledge-retrieval or generative language model.

Model notes

This is IBM's noncommercial Granite Speech 5.0 TurboCTC variant. It has approximately 470 million parameters and uses an encoder-only Conformer architecture trained with CTC. The model accepts English audio and returns transcription text. It uses SentencePiece tokenization, while the closely related Apache-licensed Granite Speech 5.0 470M TurboCTC uses BPE tokenization. The model is distributed under CC-BY-NC-SA-4.0 for research and noncommercial use only. IBM release materials report approximately 4.85% aggregate WER on public OpenASR short-form English test sets and more than 12,600 RTFx on an NVIDIA H200 under batched benchmark conditions. The model is available as downloadable weights through Hugging Face and can be used with Transformers. The model's streaming capability refers to low-latency or streaming-oriented inference and demonstrations; it is not a hosted streaming API product.

Cost

Model pricing

Input No official hosted API price; downloadable model weights
Output No official hosted API price; downloadable model weights
Model guide

Granite Speech 5.0 470M TurboCTC NC: Fast Noncommercial English Transcription

IBM Granite Speech 5.0 470M TurboCTC NC is a compact, encoder-only automatic speech recognition model for English transcription. Its 470-million-parameter Conformer and CTC architecture is designed for high-throughput, low-latency inference on local and edge hardware, while its CC-BY-NC-SA-4.0 license limits use to research and noncommercial applications.

What is Granite Speech 5.0 470M TurboCTC NC?

Granite Speech 5.0 470M TurboCTC NC is an automatic speech recognition (ASR) model from IBM Granite. Its primary job is narrow and clearly defined: it accepts English speech audio and produces a text transcription. It is not a general conversational assistant, a speech-generation system, or a multimodal language model.

The model was released on August 25, 2026, and contains approximately 470 million parameters. The NC suffix identifies the noncommercial variant. IBM distributes this checkpoint under the CC-BY-NC-SA-4.0 license, which permits research and other noncommercial use subject to the license terms. Commercial users should evaluate an appropriately licensed alternative, including the closely related Apache 2.0 Granite Speech 5.0 470M TurboCTC model where its capabilities and license meet their requirements.

The model is available as downloadable weights through the IBM Granite organization on Hugging Face rather than as a metered IBM-hosted transcription API product. Users therefore run it in their own environment or on infrastructure they operate.

Architecture and how it transcribes speech

Granite Speech 5.0 470M TurboCTC NC uses an encoder-only Conformer architecture trained with Connectionist Temporal Classification (CTC). A Conformer combines attention-based processing with convolutional processing to analyze speech patterns. CTC provides a way to align audio frames with transcript tokens without requiring a separate autoregressive language-model decoder.

In practical terms, the model processes the audio representation in a forward pass and then applies greedy CTC decoding to produce text. This is different from an autoregressive speech or language model that generates one token at a time while repeatedly consulting a decoder. The encoder-only design reduces the amount of computation needed for transcription and helps explain the model's focus on throughput and latency.

The documented architecture includes 16 Conformer blocks, chunkwise or block-based attention, self-conditioning at an intermediate layer, and several stages of temporal subsampling. These choices reduce the sequence length that later layers need to process. The model uses a SentencePiece-based output tokenizer and generates approximately 12.5 output tokens per second after temporal downsampling from the input log-mel spectrogram representation.

Capabilities and supported modalities

The supported input and output are straightforward:

  • Input: English speech audio.
  • Output: Text transcription.
  • Audio output: None.
  • Image and video input or output: None.
  • Structured output: No documented native structured-output mode.

The model is intended for English ASR. The supplied documentation does not describe multilingual transcription, speech translation, keyword biasing, speaker attribution, or diarization. It also does not include a general-purpose language-model decoder for answering questions about the transcript or conducting a spoken conversation.

The model documentation does not specify a context-window limit or maximum output-token limit. Those fields should therefore be treated as undocumented rather than assumed to be unlimited. Audio duration, memory consumption, and practical throughput will depend on the implementation, input segmentation, hardware, and batch configuration.

Speed, efficiency, and reported performance

Speed is the model's central practical advantage. IBM and the Granite Speech release materials report an aggregate word error rate of approximately 4.85% on the public short-form English test sets used for the OpenASR evaluation as of August 25, 2026. The same release materials report more than 12,600 real-time factors per second on an NVIDIA H200 under batched inference.

These are provider or release-material claims under specified benchmark conditions, not guarantees for every deployment. Word error rate can change substantially with accents, background noise, microphones, far-field recordings, conversational speech, and specialized vocabulary. The H200 result also reflects a high-end accelerator and batching. A laptop, smartphone, browser runtime, or small server will produce different results.

Compared with a larger speech-language model, this checkpoint gives up breadth in exchange for a smaller, more focused inference design. Compared with a hosted transcription service, local weights may provide more control over deployment and data handling, but the user must supply the hardware, runtime, monitoring, and maintenance.

Deployment and practical usage

Granite Speech 5.0 470M TurboCTC NC can be downloaded from Hugging Face and used with Hugging Face Transformers and compatible tooling. The model card identifies support for the Transformers automatic-speech-recognition pipeline and direct loading with AutoModelForCTC.

Because the model is a downloadable checkpoint, there is no official per-token, per-minute, or per-request hosted price associated with it. The effective cost comes from the equipment or cloud infrastructure used to run inference, plus storage, engineering, and operational costs. That can be attractive for high-volume batch transcription or environments that need local processing, but it is not the same as a zero-cost production service.

The metadata describes streaming support as low-latency or streaming-oriented inference and demonstrations. This should not be interpreted as access to an IBM hosted streaming API. Teams requiring a managed endpoint must build or select their own serving layer and verify that the chosen implementation meets their latency and audio-chunking requirements.

Reasoning, coding, and tool support

This is a specialized transcription model, so conventional language-model features are not its purpose. It does not provide a documented reasoning mode, code generation capability, function calling, tool use, web search, or agent workflow support. Any reasoning or coding scores associated with the catalog record are editorial classification fields, not provider-published capabilities or benchmark claims.

It can be part of a larger application that performs downstream reasoning or tool calls after transcription. For example, an application could use this model to transcribe a recorded meeting and then pass the resulting text to a separate language model for summarization. That would be an application pipeline, not a capability built into Granite Speech 5.0 470M TurboCTC NC itself.

License, limitations, and risk factors

The most important limitation is the license. CC-BY-NC-SA-4.0 restricts the model to research and noncommercial use. A customer-facing transcription service, paid product, revenue-generating workflow, or internal business system may require a different license depending on how it is used. Organizations should review the actual license terms and obtain legal guidance where necessary rather than assuming that local deployment makes commercial use permissible.

  • English-only according to the supplied model documentation.
  • Produces transcription text, not synthesized speech.
  • Does not provide image, video, music, or general audio generation.
  • Does not include a general-purpose language-model decoder.
  • Is not documented as a speech-translation, diarization, or speaker-attribution model.
  • Has no documented hosted API price, context limit, or maximum output-token limit.
  • Benchmark results may not transfer to noisy, accented, far-field, or domain-specific audio.

When to choose this model

Choose Granite Speech 5.0 470M TurboCTC NC when the main requirement is fast English speech-to-text transcription and the project is genuinely noncommercial. It is a strong candidate for research experiments, local transcription tools, browser or edge demonstrations, large offline batches, and prototypes where low latency matters more than conversational breadth.

Its compact 470-million-parameter size and encoder-only design are also useful when a team wants to investigate local inference rather than send audio to a hosted service. Local execution can support deployment-control requirements, although privacy and security still depend on how the surrounding application stores, transmits, and processes audio and transcripts.

Another option may be more appropriate when commercial licensing is required, when the application needs multilingual speech recognition, or when it needs speaker diarization, translation, keyword biasing, or conversational speech interaction. A larger speech-language model may be preferable when transcription is only one step in a broader audio-understanding workflow. A managed API may be preferable when the team does not want to operate model serving, scaling, and monitoring infrastructure.

Overall, Granite Speech 5.0 470M TurboCTC NC is best understood as a focused, high-throughput English ASR checkpoint rather than a general AI assistant. Its value comes from efficient local transcription and its main constraint is the noncommercial license. Those two facts should drive the deployment decision more than the model's benchmark score alone.


Answers to Frequently Asked Questions

What are the main limitations of Granite Speech 5.0 470M TurboCTC NC?
The model is documented for English speech transcription only. It does not provide audio generation, image or video processing, speech translation, speaker diarization, speaker attribution, keyword biasing, reasoning, coding, tool use, or a general conversational interface. Its benchmark results may also be less reliable for noisy, accented, far-field, conversational, or specialized-domain audio.
How can Granite Speech 5.0 470M TurboCTC NC be deployed?
The model can be downloaded from Hugging Face and run in a user-managed environment with Hugging Face Transformers and compatible tooling. The model card identifies support for the automatic-speech-recognition pipeline and direct loading with AutoModelForCTC. It is not an IBM-hosted metered transcription API, so users must provide the serving infrastructure, hardware, monitoring, and maintenance.
How fast and accurate is Granite Speech 5.0 470M TurboCTC NC?
IBM and the Granite Speech release materials report an aggregate word error rate of approximately 4.85% on the specified public short-form English OpenASR test sets and more than 12,600 real-time factors per second on an NVIDIA H200 with batched inference. Actual accuracy and speed will vary depending on audio quality, accents, noise, hardware, batching, and implementation.
What is Granite Speech 5.0 470M TurboCTC NC used for?
Granite Speech 5.0 470M TurboCTC NC is an automatic speech recognition model for converting English speech audio into text. It is designed for fast local transcription rather than conversation, speech generation, translation, diarization, or general-purpose language-model tasks.
Is Granite Speech 5.0 470M TurboCTC NC free for commercial use?
No. The NC variant is distributed under the CC-BY-NC-SA-4.0 license, which restricts it to research and other noncommercial use subject to the license terms. Organizations planning commercial, customer-facing, or revenue-generating deployments should review the license and consider an appropriately licensed alternative.


Sources 5
Provider

About IBM watsonx