Granite 4.0

Granite-4.0-H-Small

by IBM watsonx · Available; open-weight instruct model

IBM Granite-4.0-H-Small is an open-weight Apache 2.0 text model with a 128K-class context window, hybrid Mamba-2/Transformer MoE architecture, tool calling, structured output, multilingual support, and low listed watsonx.ai token pricing. It is aimed at enterprise RAG, extraction, support automation, agents, and routine coding rather than multimodal work or frontier reasoning. Maximum output tokens, knowledge cutoff, streaming, caching, and batch availability are not verified.

Text Reasoning Coding
Granite-4.0-H-Small is an open-weight language model from IBM’s Granite 4.0 family. It is designed for practical enterprise workloads such as retrieval-augmented generation (RAG), document extraction, customer-support automation, function calling, multilingual instruction following, and long-context processing. The model accepts and produces text only, but it can call tools or functions when integrated into an application. Its main attraction is the balance between a large context window, fast expected serving characteristics, and low published watsonx.ai token pricing rather than frontier-level reasoning.
Outputs

What Granite-4.0-H-Small can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Fine-tuning JSON mode Structured output
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite 4.0
Model type General Purpose
Context window 131K tokens
Release date 2025-10-02
Status Available; open-weight instruct model
Knowledge cutoff notes

IBM's official model documentation does not publish a knowledge-cutoff date for Granite-4.0-H-Small.

Model notes

Canonical Hugging Face model identifier: ibm-granite/granite-4.0-h-small. IBM watsonx.ai API identifier: ibm/granite-4-h-small. The model has 32B total parameters and approximately 9B active parameters using a hybrid Mamba-2/Transformer Mixture-of-Experts architecture with shared experts. IBM documents a 128K context window and training samples up to 512K tokens, with performance validated up to 128K. It is an instruction-tuned model derived from Granite-4.0-H-Small-Base and supports summarization, classification, extraction, question answering, RAG, code tasks, function calling, multilingual dialogue, and fill-in-the-middle code completion. Supported languages include English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. Released under Apache 2.0. Fine-tuning support is available through the open model ecosystem, but IBM does not document a separate exact-model fine-tuning endpoint. Knowledge cutoff, prompt caching, batch API availability, maximum output-token limit, and native streaming support were not directly verified for the exact model.

Cost

Model pricing

Input $0.0000636 per 1,000 tokens on IBM watsonx.ai multitenant hardware
Output $0.000265 per 1,000 tokens on IBM watsonx.ai multitenant hardware
Model guide

Granite-4.0-H-Small: IBM’s Efficient Long-Context Model for RAG and Agents

Granite-4.0-H-Small is IBM’s open-weight, text-only instruction model for long-context retrieval-augmented generation, tool-using agents, multilingual dialogue, extraction, and code tasks. Its hybrid Mamba-2/Transformer Mixture-of-Experts design combines a 128K-class context window with relatively low serving cost, although its maximum output limit, knowledge cutoff, streaming support, and exact-model fine-tuning endpoint are not documented in the supplied sources.

What is Granite-4.0-H-Small?

Granite-4.0-H-Small is an instruction-tuned language model provided by IBM. The model is available as an open-weight release under the Apache 2.0 license, with the canonical Hugging Face identifier ibm-granite/granite-4.0-h-small. IBM’s watsonx.ai identifier is ibm/granite-4-h-small.

The model belongs to IBM’s Granite 4.0 family and was released on October 2, 2025, according to the supplied model data. It is intended for applications that need a general text model with long-context handling, structured responses, code assistance, retrieval workflows, or tool calling. It is not an image, audio, or video model.

In practical terms, Granite-4.0-H-Small is aimed at developers and organizations that want an efficient model they can use through IBM watsonx.ai or obtain as an open model for their own supported deployment environment. IBM positions Granite models for business and enterprise use, while the model card provides the more specific open-weight distribution and usage information.

Architecture and position in IBM’s model lineup

IBM describes Granite-4.0-H-Small as a hybrid Mamba-2/Transformer Mixture-of-Experts model with shared experts. It has 32 billion total parameters and approximately 9 billion active parameters. The distinction matters: the total parameter count describes the full model, while the active count indicates that only part of the model is used for a given token through the mixture-of-experts design.

This architecture is intended to reduce the amount of computation required for each token compared with a dense model of the same total size. That helps explain why the model is positioned as a relatively efficient option for long-context and production workloads. The architecture alone does not guarantee a particular latency on every deployment; hardware, batching, context length, quantization, and serving software all affect performance.

Granite-4.0-H-Small is better understood as a task-oriented enterprise language model than as a consumer chatbot. Its documented uses include summarization, classification, information extraction, question answering, RAG, code tasks, multilingual dialogue, function calling, and fill-in-the-middle code completion.

Context window and input limits

The supplied model data lists a context length of 131,072 tokens, equivalent to a 128K-token context window. IBM documentation describes the model as supporting a 128K context window and notes that training samples reached up to 512K tokens, with performance validated up to 128K. The 512K training-sample figure should not be interpreted as the supported production inference limit.

A large context window is useful when an application needs to provide lengthy retrieved documents, multiple policy files, conversation history, source code, or a large collection of records in one request. It does not automatically mean that every long prompt will produce equally good answers. Retrieval quality, document ordering, repetition, and the model’s ability to identify relevant passages still affect the result.

The supplied sources do not verify a separate maximum output-token limit for Granite-4.0-H-Small. Applications should therefore set output limits according to the serving interface being used rather than assuming that the full context window is available for generated text.

Capabilities and supported modalities

Granite-4.0-H-Small is a text-input and text-output model. It does not natively accept images, audio, or video, and it does not generate those media types. Files can still be processed indirectly if an application extracts their text or uses a separate document-processing service before sending the content to the model.

CapabilitySupported or documented status
Text input and outputYes
Context window128K tokens in IBM documentation; supplied data records 131,072 tokens
Image, audio, or video inputNo native multimodal input documented
Image, audio, video, or speech outputNo
Tool or function callingSupported
Structured output and JSON modeRecorded as supported in the supplied model data
Fine-tuningSupported through the open-model ecosystem; no separate exact-model IBM endpoint was verified
Maximum output tokensNot verified
Knowledge cutoffNot published in the supplied IBM documentation

Function calling allows the model to return a structured request for an application-defined operation, such as looking up an order, querying a database, or creating a support ticket. The model does not perform those external actions by itself; the surrounding application must validate the request, execute the tool, and provide the result back to the model.

Reasoning, coding, and performance trade-offs

Granite-4.0-H-Small is suitable for ordinary business reasoning: following multi-step instructions, comparing retrieved information, extracting fields, classifying text, and deciding when to call a supplied function. The supplied editorial assessment gives it a reasoning score of 7 out of 10 and a coding score of 7 out of 10. These are comparative editorial indicators, not IBM-published benchmark results.

The model supports code generation and fill-in-the-middle completion, making it useful for code explanation, routine generation, transformation, and developer assistance. It should not automatically be treated as a frontier coding specialist. For complex software engineering, difficult mathematical reasoning, or tasks requiring extensive independent verification, a more reasoning-focused or larger model may be more appropriate.

The supplied editorial indicators rate speed at 8 out of 10 and cost at 9 out of 10. These scores reflect the model’s intended efficiency and published token prices, not a guaranteed response time or universal cost advantage. Long prompts, high concurrency, hardware selection, and output length can materially change the total cost and latency.

Pricing and access

On IBM watsonx.ai multitenant hardware, the supplied pricing data lists an input price of $0.0000636 per 1,000 tokens and an output price of $0.000265 per 1,000 tokens. These are token-based usage prices, not a monthly subscription price. A request’s total charge depends on the number of input and output tokens, and IBM may change pricing or apply product-specific conditions.

The open-weight release provides another route for use, but self-hosting is not cost-free. Infrastructure, storage, model serving, monitoring, and engineering support become the operator’s responsibility. The watsonx.ai route may be simpler for teams that want IBM-managed access, while open deployment may be more attractive when organizations need greater control over infrastructure or data handling.

Before production deployment, users should confirm the current watsonx.ai model identifier, regional availability, account requirements, throughput limits, and billing terms. The supplied sources do not verify a universal free tier or a fixed maximum output setting for this exact model.

Where Granite-4.0-H-Small fits best

  • Enterprise RAG: Use the 128K-class context window to combine retrieved policies, manuals, contracts, or support records with a question.
  • Tool-using agents: Use function calling for workflows that query business systems or trigger controlled actions.
  • Customer-support automation: Generate answers from approved knowledge sources, classify requests, and extract case details.
  • Document processing: Extract fields, summarize long text, classify records, and answer questions over converted document content.
  • Multilingual applications: IBM lists English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese among the supported languages.
  • Code assistance: Generate routine code, explain snippets, and complete missing sections where moderate coding capability is sufficient.

When to choose Granite-4.0-H-Small

Choose Granite-4.0-H-Small when the application is primarily text-based, benefits from a long context, needs tool or function calling, and places meaningful emphasis on operating cost or serving efficiency. It is particularly reasonable for teams already using IBM watsonx.ai or organizations that want an Apache 2.0 open-weight model for enterprise-oriented workflows.

Its cost and speed profile may make it preferable to a larger, slower model for high-volume extraction, classification, RAG responses, and routine support interactions. The model’s approximately 9 billion active parameters also make its design notably different from a dense model with a comparable total parameter count, although deployment results must be measured on the target hardware.

Another option may be more appropriate if the application requires image understanding, audio or video processing, media generation, native web search, exceptionally difficult reasoning, or a documented knowledge cutoff. Granite-4.0-H-Small is also not the best choice when a team requires a clearly documented exact-model fine-tuning endpoint, maximum output limit, prompt-caching policy, batch API, or streaming behavior; those details were not verified in the supplied sources.

Limitations to check before deployment

The most important limitation is modality: the model is text-only. A second is documentation uncertainty around several operational details. IBM does not publish a knowledge-cutoff date for this model in the supplied documentation, and the exact maximum output-token limit, native streaming support, prompt caching, and batch API availability were not directly verified.

Fine-tuning is available through the open-model ecosystem, but the research does not establish a dedicated IBM fine-tuning endpoint for this exact model. Organizations should also test multilingual quality, extraction accuracy, tool-call reliability, and long-context retrieval performance on their own data rather than relying only on architecture or editorial scores.

Overall, Granite-4.0-H-Small is a practical candidate for cost-conscious, long-context enterprise text workloads. Its strongest case is not universal capability; it is the combination of open-weight availability, a 128K-class context window, tool support, multilingual instruction following, and low listed watsonx.ai token prices. Those advantages should be weighed against its text-only design and the operational specifications that remain undocumented.


Answers to Frequently Asked Questions

What are the pricing and deployment options for Granite-4.0-H-Small?
On IBM watsonx.ai multitenant hardware, the supplied pricing data lists $0.0000636 per 1,000 input tokens and $0.000265 per 1,000 output tokens. The model can also be deployed from its open-weight release, although self-hosting requires infrastructure, storage, serving, monitoring, and engineering resources. Pricing, availability, throughput limits, and account requirements should be confirmed with IBM before production use.
Does Granite-4.0-H-Small support RAG and function calling?
Yes. Granite-4.0-H-Small is designed for retrieval-augmented generation and can use long prompts containing retrieved documents, policies, manuals, or records. It also supports tool or function calling, allowing an application to request operations such as database lookups or ticket creation. The surrounding application must validate and execute those tool requests.
What is Granite-4.0-H-Small?
Granite-4.0-H-Small is IBM’s instruction-tuned, open-weight language model for text-based applications such as enterprise RAG, summarization, information extraction, coding assistance, structured responses, and tool calling. It is available under the Apache 2.0 license with the Hugging Face identifier ibm-granite/granite-4.0-h-small and the watsonx.ai identifier ibm/granite-4-h-small.
How large is Granite-4.0-H-Small’s context window?
Granite-4.0-H-Small supports a 128K-token context window. The supplied model data records this as 131,072 tokens, while IBM documentation describes a 128K context window. IBM also reports training samples of up to 512K tokens, but that figure should not be treated as the supported production inference limit.


Sources 5
Provider

About IBM watsonx