Granite Speech 5.0

Granite Speech 5.0 470M TurboCTC

by IBM watsonx · Current; open-weight model

An open-weight 470-million-parameter English ASR model using conformer architecture and non-autoregressive CTC decoding for fast local and enterprise transcription.

Text Reasoning Coding
IBM Granite Speech 5.0 470M TurboCTC is a compact English speech-to-text model released on August 25, 2026. It uses a conformer acoustic encoder and non-autoregressive CTC decoding to prioritize transcription speed, making it suitable for local applications, laptops, smartphones, and enterprise workloads requiring high-throughput English ASR.
Outputs

What Granite Speech 5.0 470M TurboCTC can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
10/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite Speech 5.0
Model type Speech Recognition
Release date 2026-08-25
Status Current; open-weight model
Model notes

Granite Speech 5.0 470M TurboCTC is an approximately 470-million-parameter English ASR model using a conformer encoder, CTC training, non-autoregressive greedy decoding, 16 conformer blocks, block attention, temporal downsampling, and a 16,384-unit BPE classification head. IBM reports approximately 60,000 hours of English training data and very high benchmark throughput. The model is licensed under Apache 2.0 and is distributed as an open-weight checkpoint. Official usage is documented with Transformers, mlx-audio, and local GGUF tooling. Reported benchmark throughput depends on hardware and batching configuration.

Model guide

IBM Granite Speech 5.0 470M TurboCTC: Fast Local English Transcription

Granite Speech 5.0 470M TurboCTC is an open-weight, 470-million-parameter English automatic speech recognition model designed for very high-throughput, low-latency transcription on edge and local hardware.

What Granite Speech 5.0 470M TurboCTC is

Granite Speech 5.0 470M TurboCTC is an open-weight automatic speech recognition (ASR) model from IBM Granite. ASR models convert spoken audio into written text. This checkpoint is designed specifically for English transcription, not for general conversation, text generation, translation, speech synthesis, or multimodal reasoning.

The model has approximately 470 million parameters, making it substantially more compact than many general-purpose speech or language systems. Its design favors fast inference and efficient deployment. IBM describes it as suitable for high-throughput, low-latency transcription in enterprise applications, laptops, smartphones, and other edge environments.

It was released on August 25, 2026, and is distributed as an open-weight checkpoint under the Apache 2.0 license. The model is available through its IBM Granite model repository rather than as a separately priced, token-based hosted API product.

How the model works

Granite Speech 5.0 470M TurboCTC uses a conformer acoustic encoder. Conformers combine convolutional processing, which helps capture local audio patterns, with attention mechanisms that model relationships across a longer portion of the recording. The architecture contains 16 conformer blocks, block self-attention, self-conditioned CTC, and temporal downsampling.

Its output classification head contains 16,384 BPE units. BPE, or byte-pair encoding, is a way of representing text using reusable word and subword pieces. This allows the model to produce transcriptions without requiring a separate word-level vocabulary for every possible term.

The model is trained with Connectionist Temporal Classification (CTC). CTC is commonly used for speech recognition because it can learn the relationship between audio frames and transcribed text without requiring every individual sound to be manually aligned with a precise point in the recording.

During inference, the model uses non-autoregressive greedy decoding. In practical terms, it does not generate the transcription one token at a time in the same way as an autoregressive language model. Instead, it can process its acoustic representations more directly, reducing decoding overhead and helping it achieve low latency.

Training and verified specifications

According to the supplied IBM Granite materials, training used approximately 60,000 hours of English audio from public speech datasets and synthetic data. The model card identifies the checkpoint as an English ASR model and documents support for Transformers-based inference.

SpecificationDetails
ProviderIBM
Model familyGranite Speech 5.0
Model typeAutomatic speech recognition
Parameter countApproximately 470 million
Primary inputEnglish audio
Primary outputText transcription
ArchitectureConformer encoder with CTC-based decoding
DecodingNon-autoregressive greedy decoding
LicenseApache 2.0
Official hosted priceNot specified for this checkpoint

The reported architecture and training details are provider or project specifications. They should be distinguished from performance expectations in a particular application: actual speed and accuracy depend on the hardware, audio quality, recording conditions, batching, and integration.

Speed and performance positioning

The main reason to choose this model is throughput. IBM reports more than 12,600 times real-time throughput on an NVIDIA H200 in a reported batched OpenASR and FFASR benchmark configuration. That means the benchmark processed a very large amount of audio relative to elapsed time under the stated setup. It should not be interpreted as a guaranteed result for a laptop, smartphone, unbatched request, or production pipeline.

The model's conformer architecture, temporal downsampling, and non-autoregressive CTC decoding all support its speed-oriented design. The approximately 470-million-parameter size also makes it more practical for local deployment than larger speech systems in environments with limited compute or strict data-handling requirements.

Speed can involve trade-offs. This model is optimized for transcription rather than broad language understanding or generative flexibility. A larger or more general speech system may be more appropriate when the application needs multilingual coverage, translation, conversational interpretation, richer punctuation or formatting behavior, or integrated reasoning after transcription. The supplied research does not provide a direct accuracy comparison with a named alternative, so such comparisons should be tested on representative audio rather than inferred from the parameter count.

Supported inputs and outputs

The model accepts audio and produces text. Its primary use is English speech recognition, so an application can use it for recordings such as meetings, calls, interviews, voice notes, media files, or spoken commands that need to become searchable or processable text.

  • Audio input: Supported for English speech transcription.
  • Text output: Supported as the transcription result.
  • Image input or output: Not supported by this model.
  • Video input or output: Not supported as a native model modality.
  • Audio output: Not supported; the model does not synthesize speech.
  • Embeddings: Not provided as a documented output of this checkpoint.
  • Tool or function calling: Not supported as a model capability.
  • Structured JSON output: Not identified as a native feature.

For video workflows, the model could be used as one component after an application extracts an audio track, but that would be application-level processing rather than native video understanding. Similarly, a separate language model or application layer would be needed to summarize, classify, extract fields from, or act on the resulting transcript.

Reasoning, coding, and context limits

Granite Speech 5.0 470M TurboCTC is not a general-purpose reasoning model. It recognizes spoken language and emits text; it does not independently analyze a question, plan a task, write software, browse the web, or execute tools. Any reasoning or coding workflow would require a downstream text model or conventional software operating on the transcript.

No context-window length or maximum output-token limit is specified in the supplied research. Because this is an audio transcription checkpoint rather than a chat model, token-based context and output limits should not be assumed. Applications should follow the model's implementation documentation and test the maximum recording duration, memory use, and decoding behavior for their selected hardware and software stack.

Deployment options

IBM documents native use with the Hugging Face Transformers ecosystem, including the automatic speech recognition pipeline or direct loading with AutoModelForCTC and AutoProcessor. This makes it suitable for developers who want to manage inference in their own Python or machine-learning environment.

The model documentation also describes Apple Silicon inference through mlx-audio. Local GGUF-based inference is documented through transcribe.cpp, providing another route for compatible local deployments and quantized workflows. These options can be useful when audio cannot be sent to an external service, when predictable infrastructure costs matter, or when an application needs to operate offline or near the data source.

Deployment remains self-managed. The supplied sources do not identify an official IBM per-token endpoint or a recurring hosted price for this exact checkpoint. Infrastructure costs therefore depend on the user's hardware, cloud instance, storage, batching strategy, and operational requirements.

Best use cases

  • High-volume English meeting or interview transcription.
  • Call-processing pipelines where low latency and throughput are important.
  • Voice interfaces that need a local speech-recognition component.
  • Media transcription and searchable audio archives.
  • Edge applications running on laptops, smartphones, or other local hardware.
  • Enterprise systems with data-residency or privacy requirements that favor self-managed inference.
  • Batch processing where avoiding per-minute hosted API charges is valuable.

Its open-weight Apache 2.0 distribution may also make it attractive for teams that need to inspect, integrate, and operate the model within their own software stack. However, open-weight availability shifts responsibility for serving, scaling, monitoring, updates, and hardware compatibility to the deploying organization.

Limitations and when to choose another option

The most important limitation is language coverage: Granite Speech 5.0 470M TurboCTC is intended for English only. It is therefore a poor fit for multilingual or language-identification requirements unless the application adds other models and routing logic.

It is also an ASR model, not an end-to-end conversational assistant. Choose another type of system when the core requirement is speech translation, speech generation, audio question answering, image or video understanding, or direct tool execution. A separate language model may be needed for transcript summarization, extraction, reasoning, or coding.

Self-managed deployment is another trade-off. A hosted speech API may be more convenient for teams that do not want to operate model infrastructure, while Granite Speech 5.0 470M TurboCTC may be preferable when local control, throughput, or predictable infrastructure ownership matters more than managed convenience.

Finally, IBM's headline throughput figure comes from a specific benchmark setup. Before selecting the model, test it with the microphones, accents, background noise, audio lengths, concurrency, and hardware used by the real application. The model is best understood as a speed-focused English transcription component, not as a universal speech or language platform.

Overall assessment

Granite Speech 5.0 470M TurboCTC occupies a focused position in IBM's Granite lineup: an open-weight, compact English ASR checkpoint built for fast local and enterprise transcription. Its conformer-and-CTC design, Apache 2.0 license, and support for Transformers, Apple Silicon, and compatible local tooling make it especially relevant to developers who value deployment control and throughput.

It is a strong candidate when the requirement is straightforward English speech-to-text at scale. It is less suitable when the application needs multilingual recognition, native audio generation, general reasoning, or a managed hosted API with published per-token pricing.


Answers to Frequently Asked Questions

Is IBM Granite Speech 5.0 470M TurboCTC a general-purpose conversational or reasoning model?
No. It is focused on English speech-to-text transcription and does not natively provide reasoning, coding, summarization, translation, speech synthesis, image or video understanding, tool calling, or structured JSON generation. These capabilities require additional software or downstream language models.
How can developers run IBM Granite Speech 5.0 470M TurboCTC locally?
The model supports the Hugging Face Transformers ecosystem, including automatic speech recognition pipelines and direct use with AutoModelForCTC and AutoProcessor. IBM also documents Apple Silicon inference through mlx-audio and compatible local or quantized deployment through transcribe.cpp.
How fast is IBM Granite Speech 5.0 470M TurboCTC?
IBM reports more than 12,600 times real-time throughput on an NVIDIA H200 in a specific batched OpenASR and FFASR benchmark configuration. Actual speed depends on hardware, batching, audio quality, recording length, concurrency, and deployment configuration, so the benchmark result is not a guaranteed performance level.
What is IBM Granite Speech 5.0 470M TurboCTC used for?
IBM Granite Speech 5.0 470M TurboCTC is an open-weight automatic speech recognition model for converting English audio into text. It is designed for fast, high-throughput transcription in applications such as meetings, calls, interviews, voice interfaces, media archives, and edge deployments.
Does IBM Granite Speech 5.0 470M TurboCTC support languages other than English?
No. Granite Speech 5.0 470M TurboCTC is intended for English speech transcription. Multilingual recognition, language identification, or speech translation would require other models or additional routing logic.


Sources 4
Provider

About IBM watsonx