Granite Speech 4.1

Granite Speech 4.1 2B

by IBM watsonx · Current open-weight model

IBM Granite Speech 4.1 2B is an Apache 2.0 open-weight model for multilingual transcription and speech translation. It accepts audio and text prompts, supports keyword biasing, punctuation-aware output, and local deployment through Transformers, vLLM, llama.cpp, or MLX Audio. It returns text rather than audio and does not provide streaming transcription in the official library.

Text Reasoning Coding
Granite Speech 4.1 2B is a compact IBM Granite model built for multilingual speech-to-text and speech translation rather than general-purpose conversation. It can transcribe audio, add punctuation and capitalization, bias recognition toward important keywords, and translate speech involving English and several other languages. The Apache 2.0 model is intended for local or self-managed deployment through Transformers, vLLM, llama.cpp, and MLX Audio.
Outputs

What Granite Speech 4.1 2B can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite Speech 4.1
Model type Other
Context window 128K tokens
Release date April 29, 2026
Status Current open-weight model
Knowledge cutoff notes

No explicit knowledge-cutoff date is provided in the model card. This is a speech recognition and translation model rather than a general knowledge model.

Model notes

Open-weight Apache 2.0 model with approximately 2B parameters. Supports automatic speech recognition and speech translation for English, French, German, Spanish, Portuguese, and Japanese. Translation workflows also support English-to-Italian and English-to-Mandarin. The model supports punctuation, truecasing, and keyword list biasing. The official Granite Speech library accepts arbitrary-length audio through windowed processing, but streaming is not implemented. Granite Speech 4.1 2B Plus adds speaker-attributed transcription and word-level timestamps; Granite Speech 4.1 2B NAR is a separate higher-throughput variant. IBM documents Transformers, vLLM, llama.cpp, and MLX Audio deployment. No IBM-hosted model pricing is specified because the primary distribution is open-weight.

Model guide

Granite Speech 4.1 2B: Open-Weight Multilingual Transcription and Speech Translation

Granite Speech 4.1 2B is an open-weight IBM speech-language model for multilingual automatic speech recognition and speech translation. It accepts audio with text prompts and returns text, supporting English, French, German, Spanish, Portuguese, and Japanese speech, with additional English-to-Italian and English-to-Mandarin translation workflows.

What Granite Speech 4.1 2B is

Granite Speech 4.1 2B is an open-weight speech-language model from IBM Granite. Its main job is to turn spoken audio into text and to translate spoken language into text in another language. It is not positioned as a general chatbot, speech synthesizer, or speech-to-speech conversation system.

The model accepts audio together with text instructions or prompts, then produces text. This makes it suitable for applications such as meeting transcription, subtitle generation, multilingual call processing, and offline speech translation. IBM released the model on April 29, 2026, under the Apache 2.0 license. The model is distributed through the IBM Granite organization on Hugging Face rather than as a model with a documented IBM-hosted per-token API price.

Within the Granite Speech 4.1 family, this is the standard approximately 2-billion-parameter model. IBM also provides a Granite Speech 4.1 2B Plus variant with richer speaker-attribution and word-level timestamp features, as well as a separate 2B non-autoregressive variant intended for higher throughput. Those variants may be better choices when diarization, precise word timing, or maximum processing speed matters more than using the base model.

Supported languages and speech tasks

Granite Speech 4.1 2B supports automatic speech recognition for English, French, German, Spanish, Portuguese, and Japanese. Automatic speech recognition, or ASR, means converting spoken words into written text. Its translation workflows are centered on English and support translation between English and the six primary languages. IBM also documents English-to-Italian and English-to-Mandarin translation.

The model supports several useful transcription modes:

  • Raw transcription when the application needs a close representation of the recognized speech.
  • Punctuation and capitalization, including German noun capitalization.
  • Keyword-biased recognition for names, acronyms, technical vocabulary, and other terms that a standard recognizer might miss.
  • Speech translation when the output should be text in another supported language.

Keyword biasing is particularly useful for specialized recordings. For example, a company could provide product names, customer names, medical terms, or internal acronyms so that the recognizer gives those words additional attention. This does not make the model universally accurate in difficult audio, but it provides a practical way to adapt recognition to a known domain.

Architecture and context capacity

The model combines a conformer-based speech encoder, a speech-to-text modality adapter, and a Granite language-model decoder. A conformer is an architecture designed to model both local acoustic patterns and longer relationships in speech. IBM describes an encoder with 16 conformer blocks and connectionist temporal classification heads for character and byte-pair encoding representations.

A query-transformer projector reduces the size of the acoustic representations before passing them to the language-model component. The underlying language-model component is based on an intermediate Granite 4.0 1B base checkpoint. The documented context length is 128,000 tokens. This is a model context specification, not a promise that every recording can be processed in one uninterrupted pass: the official Granite Speech library processes long recordings through windows.

The supplied model information does not specify a maximum output-token limit. In practical use, output length will depend on the recording, selected task, runtime, and generation settings. The official library can process arbitrary-length audio through windowed processing, but this should not be confused with unlimited single-pass context.

Inputs, outputs, and deployment options

The primary input is audio, accompanied by text prompts. The expected audio setup is generally mono, 16 kHz input. The model returns text only. It does not generate audio, music, images, or video, and it does not provide built-in speech synthesis.

Granite Speech 4.1 2B is designed for self-managed deployment. IBM documents several routes:

  • Transformers: Load the model with IBM's speech-specific processor and model classes in the Hugging Face ecosystem.
  • vLLM: Serve the model through a high-performance inference runtime where the architecture and deployment configuration are supported.
  • llama.cpp: Run official GGUF weights locally, including quantized configurations that can reduce hardware requirements.
  • MLX Audio: Run the model on compatible Apple Silicon systems.

The official Granite Speech command-line library adds transcription and translation workflows, subtitle formats, JSON output, voice-activity-based segmentation, keyword biasing, and offline caching. Its current implementation uses windows for long recordings and does not provide streaming transcription. Applications that need live, token-by-token or continuously updating transcription should therefore treat the lack of implemented streaming as a significant limitation.

Accuracy and performance considerations

IBM reports a 5.33% mean word-error rate on the OpenASR leaderboard as of April 2026. This is a provider-reported benchmark claim rather than a guarantee for every recording. Word-error rate can change substantially with language, accent, background noise, microphone quality, overlapping speakers, vocabulary, and the amount of preparation provided through keyword biasing.

The model contains approximately 2 billion parameters, making it smaller than many general-purpose language models. That relatively compact size is useful for local and resource-conscious deployments, especially when quantized GGUF weights are appropriate. However, local performance depends on the selected hardware, precision, runtime, and audio workload. The supplied research does not provide a universal real-time factor or hardware requirement, so deployment teams should test their target recordings and machines rather than assuming a particular throughput.

Editorially, the model has a favorable speed-and-cost profile for an open-weight speech recognizer because it can be downloaded and run without a per-token IBM inference bill. That does not mean it is free to operate: users still pay for compute, storage, engineering, monitoring, and hosting. A managed speech service may be simpler for occasional workloads, while this model may be more economical or controllable for sustained private processing.

Reasoning, coding, and tool support

Reasoning is not the purpose of Granite Speech 4.1 2B. It can follow task-oriented text prompts such as requests for transcription, punctuation, keyword-aware recognition, or translation, but it should not be evaluated as a general reasoning model. The supplied model record gives it a low reasoning score as an editorial comparison, not as an IBM-published capability rating.

Likewise, coding is not a target capability. The model can produce text that happens to contain code if audio is transcribed, but it is not intended for code generation, software engineering, or repository analysis. The model record lists no tool or function-calling support. Any integration with storage, subtitle publishing, databases, or downstream business systems must be implemented by the surrounding application.

Pricing and licensing

Granite Speech 4.1 2B is released under the Apache 2.0 license and is available as an open-weight model. No IBM-hosted per-token model price is specified in the supplied research. The practical cost depends on where and how the model is run: local hardware, a rented GPU or CPU instance, an enterprise cluster, or another hosting environment.

This pricing model differs from a hosted speech API. Self-hosting can provide more control over data handling, deployment location, and ongoing usage costs, but it requires users to manage infrastructure and model serving. The Apache 2.0 license may be attractive for commercial and private deployments, subject to the license terms and the obligations of any other software or data used in the application.

Important limitations and alternatives

The base model does not include the richer speaker-attribution and word-level timestamp features offered by Granite Speech 4.1 2B Plus. If an application must identify who spoke each segment or align every word with precise timestamps, the Plus variant is the more suitable option within the same family.

The 2B non-autoregressive variant is another relevant alternative when throughput is the main concern. The supplied research identifies it as a separate higher-throughput variant, but does not provide enough detail to conclude that it is better for every language or recording condition.

Granite Speech 4.1 2B is also a poor fit for several classes of work:

  • Live transcription that requires implemented streaming support.
  • Speech synthesis or spoken responses.
  • Speech-to-speech conversational assistants.
  • Music, image, or video generation.
  • General-purpose question answering and complex reasoning.
  • Applications that require built-in tools, function calling, or external web access.

For these needs, a dedicated managed speech service, a speech-to-speech system, a general multimodal model, or a separate orchestration layer may be more appropriate. Those alternatives may offer easier operations or richer real-time features, while Granite Speech 4.1 2B offers the advantages of open weights, local execution, and domain-specific transcription controls.

When to choose Granite Speech 4.1 2B

Choose Granite Speech 4.1 2B when the central requirement is multilingual audio transcription or speech translation and you want to control the deployment environment. It is especially suitable for private or offline processing, enterprise speech-to-text pipelines, subtitles and captions, and workflows involving English, French, German, Spanish, Portuguese, or Japanese.

It is also a strong candidate when keyword biasing matters. A specialized vocabulary list can help transcription systems handle names, abbreviations, and technical terms that occur repeatedly in a particular domain. Quantized GGUF deployment may make the model useful in edge or resource-conscious environments, provided testing confirms adequate accuracy and speed.

Choose a different option when the application needs real-time streaming, speaker labels, word-level timestamps, audio output, or general conversational intelligence. Within IBM's related lineup, Granite Speech 4.1 2B Plus is the better-supported choice for speaker attribution and word timing, while the 2B NAR variant is the more directly relevant comparison for higher-throughput processing.

Bottom line

Granite Speech 4.1 2B is a focused open-weight model for turning multilingual speech into text and translating speech into text. Its useful combination of six primary speech-recognition languages, English-centered translation, keyword biasing, punctuation support, local runtimes, and Apache 2.0 licensing makes it practical for teams that want to build and operate their own transcription pipeline.

Its boundaries are equally important: it returns text rather than audio, does not provide implemented streaming in the official library, has no documented IBM-hosted usage price, and is not a general reasoning, coding, or tool-using model. Its best role is as a self-managed speech recognition and translation component, not as a complete voice assistant.


Answers to Frequently Asked Questions

What license does Granite Speech 4.1 2B use, and does it have an IBM-hosted API price?
Granite Speech 4.1 2B is released under the Apache 2.0 license as an open-weight model. The supplied information does not specify an IBM-hosted per-token API price, so operating costs depend on the hardware, hosting, storage, and engineering resources used to run it.
How can Granite Speech 4.1 2B be deployed?
Granite Speech 4.1 2B can be self-hosted through Transformers, vLLM, llama.cpp with official GGUF weights, or MLX Audio on compatible Apple Silicon systems. It is also supported by the official Granite Speech command-line library for transcription and translation workflows.
Can Granite Speech 4.1 2B perform real-time streaming transcription?
The official Granite Speech library processes long recordings through windows but does not currently provide implemented streaming transcription. Applications requiring live, continuously updating or token-by-token transcription should use a different solution or add a separate streaming system.
Which languages does Granite Speech 4.1 2B support?
The model supports automatic speech recognition for English, French, German, Spanish, Portuguese, and Japanese. Its translation workflows focus on English and the six primary languages, with additional documented support for English-to-Italian and English-to-Mandarin translation.
What is Granite Speech 4.1 2B used for?
Granite Speech 4.1 2B is an open-weight speech-language model for converting spoken audio into text and translating speech into text in another language. Typical uses include meeting transcription, subtitle generation, multilingual call processing, and offline speech translation.


Sources 4
Provider

About IBM watsonx