Granite Speech 4.1

Granite-Speech-4.1-2B-Plus

by IBM watsonx · Current open-weight model

An open-weight IBM speech-language model for English, French, German, Spanish, and Portuguese transcription. The Plus variant adds speaker-attributed ASR, word-level timestamps, keyword list biasing, and incremental decoding. It runs through self-managed tooling under an Apache 2.0 license and requires application-side parsing and timestamp rollover handling for rich transcripts.

Text Reasoning Coding
Granite-Speech-4.1-2B-Plus is an IBM Granite model built for speech-to-text rather than general-purpose chat. It accepts audio together with text instructions and can produce ordinary transcripts, speaker-labeled turns, or word-level timing tags. The model supports English, French, German, Spanish, and Portuguese, is distributed in BF16 under the Apache 2.0 license, and is intended for local or self-managed deployment. Its Plus features are particularly useful when a transcript needs speaker identity, timing metadata, or improved recognition of names and specialized terminology.
Outputs

What Granite-Speech-4.1-2B-Plus can produce

Text
Inputs

What it can understand

Text Audio Multimodal input
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite Speech 4.1
Model type Other
Context window 4K tokens
Knowledge cutoff April 2024
Release date April 28, 2026
Status Current open-weight model
Knowledge cutoff notes

The official model-card usage example defines the system prompt with a knowledge cutoff date of April 2024. Because this is a speech transcription model, the cutoff is mainly relevant to its inherited language-model component rather than to audio-recognition freshness.

Model notes

Canonical Hugging Face identifier: ibm-granite/granite-speech-4.1-2b-plus. The model supports plain ASR, speaker-attributed ASR, word-level timestamp generation, incremental decoding through prefix_text, and keyword list biasing. It supports English, French, German, Spanish, and Portuguese. Rich-transcription modes do not provide punctuation and capitalization like the base Granite-Speech-4.1-2B model. Timestamp values are centiseconds modulo 1,000, so applications must handle rollover every 10 seconds. IBM documents evaluation on audio segments up to approximately 9 minutes for ASR and speaker-attributed ASR and up to approximately 3.5 minutes for timestamps. The model is distributed in BF16 with approximately 2B parameters under Apache 2.0. It is self-hosted/open-weight and has no official IBM token pricing. The model card includes a system prompt stating a knowledge cutoff of April 2024; this is not a conventional web-search or general language-model knowledge interface.

Model guide

Granite-Speech-4.1-2B-Plus: Open-Weight Transcription with Speaker Labels and Timestamps

Granite-Speech-4.1-2B-Plus is IBM’s open-weight, 2-billion-parameter speech-language model for multilingual automatic speech recognition. Its distinguishing features are speaker-attributed transcripts, word-level timestamps, keyword list biasing, and incremental decoding, making it suitable for self-hosted meeting, interview, call, and media transcription.

What is Granite-Speech-4.1-2B-Plus?

Granite-Speech-4.1-2B-Plus is an open-weight speech-language model from IBM’s Granite family. Its primary job is automatic speech recognition (ASR): converting spoken audio into written text. It is not positioned as a general-purpose conversational model, coding assistant, image model, or speech-synthesis system.

The model combines a speech encoder with a language-model component and multimodal projection layers that connect audio features to text generation. The configuration identifies approximately 2 billion parameters, BF16 distribution, and a 4,096-token text context window. IBM makes the model available through the Granite model collection and the official IBM Granite Hugging Face organization under the identifier ibm-granite/granite-speech-4.1-2b-plus.

The model is an expanded version of Granite-Speech-4.1-2B. The important distinction is that the Plus variant adds richer transcription modes, including speaker attribution and word-level timestamps, alongside keyword biasing and incremental decoding.

What can the model transcribe?

Standard speech-to-text

In its ordinary ASR mode, Granite-Speech-4.1-2B-Plus converts an audio recording into a text transcript. The documented language coverage is English, French, German, Spanish, and Portuguese. This makes it suitable for multilingual enterprise recordings, although the supplied research does not establish equal accuracy across all five languages or across different accents, recording conditions, and domains.

Applications provide audio along with a textual instruction. The model’s output is text, so a typical integration can store the result in a transcript database, display it in an editor, or pass it to downstream search and analysis software.

Speaker-attributed transcription

Speaker-attributed ASR adds labels such as [Speaker 1]: and [Speaker 2]: before speaker turns. Speaker numbers follow the order in which people first appear in the recording. This is useful for meetings, interviews, customer calls, hearings, and other recordings where knowing who spoke is as important as knowing what was said.

This feature is designed to provide speaker-labeled output within the transcription workflow, reducing the need for a separate basic diarization step for supported use cases. The labels are returned as text tags, however, so applications should parse and validate them rather than treating them as a fully independent speaker-management database.

Word-level timestamps

Timestamp mode places a timing tag after each word using a format such as [T:N]. Values are expressed in centiseconds and are represented modulo 1,000. In practical terms, timestamp values roll over every 10 seconds, so an application processing a longer recording must unwrap the values to reconstruct continuous time.

Word timing can support subtitles, transcript alignment, searchable audio, clip selection, and synchronized review interfaces. IBM’s documented evaluations cover timestamping on shorter segments than ordinary ASR, with timestamp tests extending to approximately 3.5 minutes. This makes segmentation and timestamp rollover handling important implementation details for longer recordings.

Keyword list biasing

Keyword list biasing lets an application provide terms that deserve special attention during recognition. Examples include customer names, product names, acronyms, medical or legal terminology, and internal project names. This is especially useful when ordinary speech recognition may confuse uncommon words with more frequent alternatives.

Keyword biasing is prompt-based rather than a separate fine-tuned model in the supplied documentation. Results should therefore be evaluated on representative audio, particularly when a keyword list is long or contains similar-sounding terms.

Incremental decoding with prefix text

For segmented or continuing audio, applications can pass earlier transcript material through the prefix_text field. This can preserve accumulated context and speaker numbering while avoiding unnecessary re-decoding of text that has already been processed.

Incremental decoding is useful for long recordings and streaming-like pipelines built from successive audio segments. The research does not verify a separate hosted real-time API or a guaranteed latency target, so developers should distinguish this segmented workflow from a provider-operated live transcription service.

Technical specifications and supported modalities

SpecificationDetails
ProviderIBM
Model familyGranite Speech 4.1
ParametersApproximately 2 billion
FormatBF16
Text context4,096 tokens
InputAudio and text instructions
OutputText transcripts, speaker tags, or timestamp tags
Documented languagesEnglish, French, German, Spanish, and Portuguese
LicenseApache 2.0

Granite-Speech-4.1-2B-Plus has multimodal input because it processes audio together with text. Its direct output is text only: it does not generate speech, music, images, or video. The research does not specify a maximum output-token limit, a separate audio-duration maximum for production use, or an official streaming guarantee. IBM’s documented evaluation notes cover approximately nine minutes for ordinary and speaker-attributed ASR and approximately 3.5 minutes for timestamp generation; these figures should not automatically be treated as hard product limits.

Formatting and post-processing considerations

The rich-transcription modes return structured text markers rather than a polished document. Speaker labels and timing tags are valuable metadata, but they require application-side parsing. A production pipeline may need to convert the tags into structured records containing speaker, word, start time, and end time fields.

Unlike the base Granite-Speech-4.1-2B model, the Plus variant’s rich-transcription modes do not provide punctuation and capitalization in the same way. Applications that need publication-ready prose, captions, or readable meeting notes may therefore need a separate formatting step. That step should preserve the original words and timing information rather than accidentally changing the transcript.

Timestamp rollover is another important engineering issue. Because the centisecond tag is modulo 1,000, a sequence can appear to move backward after each 10-second interval. Software must detect the rollover and add the appropriate offset before presenting continuous timestamps to users.

Deployment, license, and pricing

IBM documents local loading with Transformers and serving through vLLM. The model card indicates that current Transformers support may be needed, and some environments may require a recent or source installation before Plus-specific functionality is available. Operators should test the exact model revision and library versions used in deployment.

The Apache 2.0 license permits research and commercial use, subject to the license terms. Open weights do not mean that operation is free: the deploying organization remains responsible for GPU or other inference infrastructure, storage, audio processing, monitoring, security, and compliance.

No official IBM token price is supplied for this model. It is presented as a self-hosted open-weight model rather than a metered IBM-hosted API model with published per-token pricing. Its effective cost depends on hardware, utilization, recording volume, engineering effort, and operational requirements. This can be attractive for organizations that already operate inference infrastructure or need local control over audio data, but less convenient for users seeking a simple pay-as-you-go transcription endpoint.

Strengths and limitations

Main strengths

  • Rich transcription: speaker labels and word-level timestamps extend the model beyond plain ASR.
  • Domain vocabulary support: keyword list biasing can help with names, acronyms, products, and technical terminology.
  • Multilingual coverage: the documented languages are English, French, German, Spanish, and Portuguese.
  • Local deployment: open weights and the Apache 2.0 license support self-managed research and commercial applications.
  • Incremental processing: prefix_text can help maintain context and speaker numbering across audio segments.

Important limitations

  • It is specialized for transcription and should not be selected as a general chat, reasoning, coding, image, video, or speech-generation model.
  • Rich output is tagged text, not a finished structured transcript database. Parsing and validation are required.
  • Timestamp values roll over every 10 seconds and must be corrected by application logic.
  • Rich modes may need downstream punctuation and capitalization restoration.
  • The supplied research does not establish a hosted API, guaranteed real-time latency, fine-tuning support, tool calling, JSON mode, or a maximum output-token limit.
  • Self-hosting transfers infrastructure, scaling, monitoring, and compliance responsibilities to the operator.

When should you choose Granite-Speech-4.1-2B-Plus?

Choose this model when the main problem is multilingual speech transcription and the output needs more than plain text. It is a strong fit for meeting records, interviews, call analytics, searchable audio archives, speaker-labeled notes, subtitle preparation, and media workflows that need word-level alignment. It is particularly relevant when audio should remain within an organization’s controlled deployment environment.

Its cost profile can also be favorable compared with a hosted transcription service at high volume, but only when the organization can use its infrastructure efficiently. The model is smaller and more specialized than large general-purpose multimodal models, which can make local inference more practical, but the supplied research does not provide benchmark measurements for speed or accuracy. The model’s recorded editorial speed score of 7 and cost score of 9 are comparative assessments, not IBM-published benchmarks.

Another speech-recognition option may be more appropriate when a project requires a turnkey hosted API, guaranteed live-streaming latency, automatic punctuation and formatting, a larger language list, or managed scaling. A general-purpose language or multimodal model may be preferable for open-ended reasoning, coding, tool use, or document generation after transcription. Conversely, a simpler ASR model may be preferable when only plain transcripts are needed and speaker labels, timestamps, or keyword biasing would add unnecessary processing complexity.

Bottom line

Granite-Speech-4.1-2B-Plus is best understood as a self-hosted transcription engine with richer metadata capabilities. Its defining advantages are speaker-attributed output, word timing, keyword biasing, multilingual support, and incremental decoding. Its trade-offs are equally clear: no published IBM token pricing, no verified hosted-service guarantees, limited output formats, timestamp rollover requirements, and a need for downstream formatting. For teams that need controlled deployment and structured speech transcripts, those trade-offs may be worthwhile; for users seeking a ready-made transcription API or a general AI assistant, another type of option is likely a better fit.


Answers to Frequently Asked Questions

How can Granite-Speech-4.1-2B-Plus be deployed, and what does it cost?
The model can be loaded locally with Transformers and served through vLLM. It is distributed under the Apache 2.0 license and has no published IBM token price because it is presented as a self-hosted open-weight model. Operators must provide and pay for the required inference infrastructure, storage, monitoring, security, and compliance.
Which languages does Granite-Speech-4.1-2B-Plus support?
The documented languages are English, French, German, Spanish, and Portuguese. Actual accuracy may vary depending on accents, recording conditions, audio quality, and the subject domain.
Does Granite-Speech-4.1-2B-Plus support speaker labels and timestamps?
Yes. It can add speaker labels such as [Speaker 1] and [Speaker 2], as well as word-level timestamp tags in the format [T:N]. Timestamp values are expressed in centiseconds modulo 1,000, so applications must handle rollover every 10 seconds when processing longer recordings.
What is Granite-Speech-4.1-2B-Plus used for?
Granite-Speech-4.1-2B-Plus is an open-weight automatic speech recognition model for converting audio into text. It is designed for multilingual transcription, speaker-attributed transcripts, word-level timestamps, keyword biasing, and incremental decoding rather than general chat, coding, image generation, or speech synthesis.


Sources 3
Provider

About IBM watsonx