What is Granite-4.0-H-Small?
Granite-4.0-H-Small is an instruction-tuned language model provided by IBM. The model is available as an open-weight release under the Apache 2.0 license, with the canonical Hugging Face identifier ibm-granite/granite-4.0-h-small. IBM’s watsonx.ai identifier is ibm/granite-4-h-small.
The model belongs to IBM’s Granite 4.0 family and was released on October 2, 2025, according to the supplied model data. It is intended for applications that need a general text model with long-context handling, structured responses, code assistance, retrieval workflows, or tool calling. It is not an image, audio, or video model.
In practical terms, Granite-4.0-H-Small is aimed at developers and organizations that want an efficient model they can use through IBM watsonx.ai or obtain as an open model for their own supported deployment environment. IBM positions Granite models for business and enterprise use, while the model card provides the more specific open-weight distribution and usage information.
Architecture and position in IBM’s model lineup
IBM describes Granite-4.0-H-Small as a hybrid Mamba-2/Transformer Mixture-of-Experts model with shared experts. It has 32 billion total parameters and approximately 9 billion active parameters. The distinction matters: the total parameter count describes the full model, while the active count indicates that only part of the model is used for a given token through the mixture-of-experts design.
This architecture is intended to reduce the amount of computation required for each token compared with a dense model of the same total size. That helps explain why the model is positioned as a relatively efficient option for long-context and production workloads. The architecture alone does not guarantee a particular latency on every deployment; hardware, batching, context length, quantization, and serving software all affect performance.
Granite-4.0-H-Small is better understood as a task-oriented enterprise language model than as a consumer chatbot. Its documented uses include summarization, classification, information extraction, question answering, RAG, code tasks, multilingual dialogue, function calling, and fill-in-the-middle code completion.
Context window and input limits
The supplied model data lists a context length of 131,072 tokens, equivalent to a 128K-token context window. IBM documentation describes the model as supporting a 128K context window and notes that training samples reached up to 512K tokens, with performance validated up to 128K. The 512K training-sample figure should not be interpreted as the supported production inference limit.
A large context window is useful when an application needs to provide lengthy retrieved documents, multiple policy files, conversation history, source code, or a large collection of records in one request. It does not automatically mean that every long prompt will produce equally good answers. Retrieval quality, document ordering, repetition, and the model’s ability to identify relevant passages still affect the result.
The supplied sources do not verify a separate maximum output-token limit for Granite-4.0-H-Small. Applications should therefore set output limits according to the serving interface being used rather than assuming that the full context window is available for generated text.
Capabilities and supported modalities
Granite-4.0-H-Small is a text-input and text-output model. It does not natively accept images, audio, or video, and it does not generate those media types. Files can still be processed indirectly if an application extracts their text or uses a separate document-processing service before sending the content to the model.
| Capability | Supported or documented status |
|---|---|
| Text input and output | Yes |
| Context window | 128K tokens in IBM documentation; supplied data records 131,072 tokens |
| Image, audio, or video input | No native multimodal input documented |
| Image, audio, video, or speech output | No |
| Tool or function calling | Supported |
| Structured output and JSON mode | Recorded as supported in the supplied model data |
| Fine-tuning | Supported through the open-model ecosystem; no separate exact-model IBM endpoint was verified |
| Maximum output tokens | Not verified |
| Knowledge cutoff | Not published in the supplied IBM documentation |
Function calling allows the model to return a structured request for an application-defined operation, such as looking up an order, querying a database, or creating a support ticket. The model does not perform those external actions by itself; the surrounding application must validate the request, execute the tool, and provide the result back to the model.
Reasoning, coding, and performance trade-offs
Granite-4.0-H-Small is suitable for ordinary business reasoning: following multi-step instructions, comparing retrieved information, extracting fields, classifying text, and deciding when to call a supplied function. The supplied editorial assessment gives it a reasoning score of 7 out of 10 and a coding score of 7 out of 10. These are comparative editorial indicators, not IBM-published benchmark results.
The model supports code generation and fill-in-the-middle completion, making it useful for code explanation, routine generation, transformation, and developer assistance. It should not automatically be treated as a frontier coding specialist. For complex software engineering, difficult mathematical reasoning, or tasks requiring extensive independent verification, a more reasoning-focused or larger model may be more appropriate.
The supplied editorial indicators rate speed at 8 out of 10 and cost at 9 out of 10. These scores reflect the model’s intended efficiency and published token prices, not a guaranteed response time or universal cost advantage. Long prompts, high concurrency, hardware selection, and output length can materially change the total cost and latency.
Pricing and access
On IBM watsonx.ai multitenant hardware, the supplied pricing data lists an input price of $0.0000636 per 1,000 tokens and an output price of $0.000265 per 1,000 tokens. These are token-based usage prices, not a monthly subscription price. A request’s total charge depends on the number of input and output tokens, and IBM may change pricing or apply product-specific conditions.
The open-weight release provides another route for use, but self-hosting is not cost-free. Infrastructure, storage, model serving, monitoring, and engineering support become the operator’s responsibility. The watsonx.ai route may be simpler for teams that want IBM-managed access, while open deployment may be more attractive when organizations need greater control over infrastructure or data handling.
Before production deployment, users should confirm the current watsonx.ai model identifier, regional availability, account requirements, throughput limits, and billing terms. The supplied sources do not verify a universal free tier or a fixed maximum output setting for this exact model.
Where Granite-4.0-H-Small fits best
- Enterprise RAG: Use the 128K-class context window to combine retrieved policies, manuals, contracts, or support records with a question.
- Tool-using agents: Use function calling for workflows that query business systems or trigger controlled actions.
- Customer-support automation: Generate answers from approved knowledge sources, classify requests, and extract case details.
- Document processing: Extract fields, summarize long text, classify records, and answer questions over converted document content.
- Multilingual applications: IBM lists English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese among the supported languages.
- Code assistance: Generate routine code, explain snippets, and complete missing sections where moderate coding capability is sufficient.
When to choose Granite-4.0-H-Small
Choose Granite-4.0-H-Small when the application is primarily text-based, benefits from a long context, needs tool or function calling, and places meaningful emphasis on operating cost or serving efficiency. It is particularly reasonable for teams already using IBM watsonx.ai or organizations that want an Apache 2.0 open-weight model for enterprise-oriented workflows.
Its cost and speed profile may make it preferable to a larger, slower model for high-volume extraction, classification, RAG responses, and routine support interactions. The model’s approximately 9 billion active parameters also make its design notably different from a dense model with a comparable total parameter count, although deployment results must be measured on the target hardware.
Another option may be more appropriate if the application requires image understanding, audio or video processing, media generation, native web search, exceptionally difficult reasoning, or a documented knowledge cutoff. Granite-4.0-H-Small is also not the best choice when a team requires a clearly documented exact-model fine-tuning endpoint, maximum output limit, prompt-caching policy, batch API, or streaming behavior; those details were not verified in the supplied sources.
Limitations to check before deployment
The most important limitation is modality: the model is text-only. A second is documentation uncertainty around several operational details. IBM does not publish a knowledge-cutoff date for this model in the supplied documentation, and the exact maximum output-token limit, native streaming support, prompt caching, and batch API availability were not directly verified.
Fine-tuning is available through the open-model ecosystem, but the research does not establish a dedicated IBM fine-tuning endpoint for this exact model. Organizations should also test multilingual quality, extraction accuracy, tool-call reliability, and long-context retrieval performance on their own data rather than relying only on architecture or editorial scores.
Overall, Granite-4.0-H-Small is a practical candidate for cost-conscious, long-context enterprise text workloads. Its strongest case is not universal capability; it is the combination of open-weight availability, a 128K-class context window, tool support, multilingual instruction following, and low listed watsonx.ai token prices. Those advantages should be weighed against its text-only design and the operational specifications that remain undocumented.

