Granite 4.0

Granite-4.0-H-Micro

by IBM watsonx · Current open-weight instruct model

Granite-4.0-H-Micro is IBM’s 3B-parameter Apache 2.0 instruct model for efficient local text generation. It offers a 128K-token context window, multilingual support, coding and fill-in-the-middle capabilities, RAG suitability, and function calling. The model is optimized for speed, cost efficiency, and deployment control, but it is not multimodal, has no published hosted token price or maximum output limit, and may be less capable than larger models on difficult reasoning and complex coding tasks.

Text Reasoning Coding
IBM Granite-4.0-H-Micro is a compact open-weight language model built for applications that need fast, affordable, and locally managed text generation rather than frontier-scale reasoning. Its hybrid architecture combines attention layers with Mamba2 layers and supports a documented 128K-token context window. The instruct model is suited to RAG, extraction, summarization, coding assistance, multilingual dialogue, and lightweight agents that call external tools.
Outputs

What Granite-4.0-H-Micro can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Fine-tuning
Model profile

Performance characteristics

5/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite 4.0
Model type Lightweight
Context window 128K tokens
Release date 2025-10-02
Status Current open-weight instruct model
Knowledge cutoff notes

IBM's official model documentation and model card do not state a specific knowledge cutoff for Granite-4.0-H-Micro.

Model notes

Canonical model identifier: ibm-granite/granite-4.0-h-micro. This is the instruct variant derived from Granite-4.0-H-Micro-Base. It uses a 3B-parameter dense hybrid architecture with four attention layers and 36 Mamba2 layers, and has a documented 128K-token sequence length. The model is open-weight and Apache 2.0 licensed; IBM does not publish an official token-based hosted API price for this exact model. The model card documents text generation, multilingual use, fill-in-the-middle code completion, retrieval-augmented generation, and function calling. Editorial scores are comparative estimates rather than IBM-provided ratings.

Model guide

IBM Granite-4.0-H-Micro: A Compact 128K-Context Model for Local AI

Granite-4.0-H-Micro is IBM’s 3-billion-parameter, Apache 2.0-licensed instruct model for low-latency local inference, long-context text generation, coding assistance, retrieval-augmented generation, multilingual applications, and tool-calling workflows.

What is Granite-4.0-H-Micro?

Granite-4.0-H-Micro is a 3-billion-parameter instruct language model from IBM’s Granite family. IBM released it on October 2, 2025, and distributes the open-weight model through its official Hugging Face organization under the Apache 2.0 license. “Instruct” means that the model has been adapted to follow user directions, answer questions, and participate in structured workflows rather than serving only as a base text predictor.

The model is designed for practical deployment where latency, hardware requirements, and operating cost matter. It is not positioned as IBM’s answer to the largest frontier models. Instead, Granite-4.0-H-Micro targets focused tasks such as summarization, classification, information extraction, question answering, retrieval-augmented generation (RAG), code assistance, and function calling.

Its open-weight distribution also makes it different from a typical hosted chatbot model. Developers can manage the inference environment themselves using compatible tools and hardware, although actual performance depends on the deployment configuration, precision, quantization, and available compute.

Where it fits in IBM’s current model lineup

Granite-4.0-H-Micro belongs to IBM’s Granite 4.0 family and is the instruct variant associated with Granite-4.0-H-Micro-Base. The “Micro” designation reflects its relatively small 3B parameter count. That size is useful when an application needs a model that can run with lower resource requirements than much larger language models.

IBM’s broader watsonx portfolio can provide managed development, governance, deployment, and model access, but Granite-4.0-H-Micro itself is documented as an open-weight model intended for self-hosted or locally managed use. IBM does not publish a token-based hosted API price for this exact model in the supplied documentation. Therefore, the cost question is primarily about infrastructure and operations rather than a fixed per-token IBM endpoint price.

Architecture and 128K context window

Granite-4.0-H-Micro uses a decoder-only dense hybrid architecture. It combines four attention layers with 36 Mamba2 layers. Attention is widely used to relate tokens across a sequence, while Mamba2 is a state-space architecture intended to process sequences efficiently. The combination is designed to balance long-context handling and inference efficiency, although real-world results will vary by software stack and hardware.

The model has a 2,048-dimensional embedding size, 32 attention heads, eight key-value heads, SwiGLU activation, RMSNorm, and shared input/output embeddings. These are verified architectural details from the model documentation rather than general estimates.

Its documented sequence length is 128,000 tokens. This is a substantial context capacity for a compact model and can be useful when working with long documents, collections of retrieved passages, code files, or extended conversations. A context window is not the same as a guaranteed output length: IBM’s supplied documentation does not state a maximum output-token limit for this model. Applications should therefore confirm the limit imposed by the selected inference framework and generation configuration.

Capabilities and supported modalities

Granite-4.0-H-Micro accepts text and produces text. It does not natively generate images, audio, video, embeddings, or other non-text media. It should consequently be evaluated as a text language model, not as a multimodal assistant or media-generation system.

The model card documents support for English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. Multilingual support does not mean equal quality across all languages; the supplied research does not establish that performance is uniform, so language-specific testing remains important.

Documented use cases include:

  • Summarizing reports, messages, and retrieved documents.
  • Classifying text and extracting structured information from unstructured content.
  • Question answering over application data or retrieved passages.
  • Generating text for multilingual assistants and business workflows.
  • Code-related tasks, including code completion and fill-in-the-middle completion.
  • RAG systems that provide relevant external text to the model at generation time.
  • Lightweight agent workflows that require function or tool calls.

Tool calling and structured workflows

The instruct variant supports tool calling through structured function definitions and tool-call messages. This allows an application to expose operations such as database lookups, calculations, ticket creation, or document retrieval. The model can select or request a tool, while the surrounding application remains responsible for executing it and returning the result.

This capability makes Granite-4.0-H-Micro a candidate for compact agents and workflow automation, particularly where the model needs to route requests or invoke a limited set of known actions. Tool calling should not be confused with autonomous access to the web or external systems: the model does not independently provide web search, and any external action must be implemented by the host application.

The supplied model research does not establish a separate provider-native JSON mode or constrained JSON-schema output feature. Developers can still request structured text or parse tool-call messages, but they should not assume that every generated response will conform perfectly to an arbitrary JSON schema without application-side validation.

Reasoning, coding, speed, and cost trade-offs

Granite-4.0-H-Micro is better understood as a compact general-purpose instruct model than as a specialist reasoning model. Its 3B size can be an advantage for routine classification, extraction, summarization, and straightforward question answering, but it may be less reliable on difficult multi-step reasoning, broad factual tasks, or complex planning than larger models. The supplied research does not provide a standardized reasoning benchmark for this specific model.

It supports coding assistance and fill-in-the-middle completion, making it useful for code suggestions, small transformations, documentation, and lightweight development tools. Larger coding models may be more appropriate for large repositories, complicated refactoring, intricate debugging, or tasks requiring extensive project-wide context and strong reasoning.

The editorial data rates its relative reasoning capability at 5 out of 10, coding at 6 out of 10, speed at 8 out of 10, and cost efficiency at 9 out of 10. These are comparative editorial estimates, not IBM-published benchmark scores. They express the model’s likely positioning: faster and cheaper to operate than larger models, while accepting lower capability on demanding tasks.

Because IBM does not publish an official hosted token price for this exact model, there is no verified per-input-token or per-output-token price to quote. For self-hosted use, total cost depends on hardware, hosting, electricity, storage, engineering effort, and the chosen inference framework. Quantization may reduce memory requirements, but the research does not specify a universal hardware minimum or a guaranteed performance level.

Deployment options and limitations

The model is intended for deployment with tools such as Transformers, vLLM, SGLang, Docker, and compatible local-inference or quantization systems. These tools provide possible deployment routes rather than a guarantee that every configuration offers identical features or performance.

Local deployment can help organizations retain control over data and avoid dependence on a particular hosted endpoint. It can also make predictable low-latency inference possible for focused workloads. The trade-off is operational responsibility: teams must select hardware, configure serving, monitor resource use, apply security controls, and maintain the model-serving environment.

Granite-4.0-H-Micro has several important limitations. It is text-only, does not provide built-in web search, and has no documented maximum output-token limit in the supplied research. Its smaller parameter count may reduce performance on difficult reasoning, extensive world knowledge, and complex coding tasks. Multilingual quality may vary by language. Like other language models, it can produce inaccurate, biased, or unsafe responses, so production applications need validation and appropriate safeguards.

When to choose Granite-4.0-H-Micro

Choose Granite-4.0-H-Micro when the main requirement is an open-weight text model that can be deployed locally or under your own infrastructure controls. It is particularly well suited to:

  • Low-latency summarization, classification, and extraction.
  • RAG systems with long retrieved context or long documents.
  • Multilingual internal assistants where text output is sufficient.
  • Small coding assistants and fill-in-the-middle completion tools.
  • Lightweight tool-calling agents with a defined set of functions.
  • Organizations seeking Apache 2.0 licensing and deployment flexibility.

A larger model may be more appropriate when the application depends on difficult reasoning, advanced coding, complex planning, or consistently strong performance across broad domains. A hosted model may be preferable when the team does not want to operate inference infrastructure. A multimodal model is required for image, audio, or video input and output, since Granite-4.0-H-Micro is limited to text.

The practical choice is therefore a trade-off rather than a claim that the model is universally superior. Granite-4.0-H-Micro prioritizes compact deployment, speed, and operating efficiency. It is most compelling when those factors and data-control requirements outweigh the benefits of a larger model’s potentially stronger reasoning or generation quality.

Bottom line

Granite-4.0-H-Micro is a compact IBM open-weight model for text-focused applications that need a large context window without adopting a much larger model. Its 3B parameter count, hybrid attention-Mamba2 design, 128K-token sequence length, multilingual support, coding features, RAG suitability, and tool-calling capability give it a practical role in local and embedded AI systems.

Its strengths are most relevant to efficient, controlled workloads rather than frontier-level reasoning. Prospective users should test the exact languages, prompts, tools, quantization settings, and hardware they plan to use, and should treat the absence of a published hosted price or maximum output limit as a deployment detail requiring further configuration-specific verification.


Answers to Frequently Asked Questions

What are the main limitations of Granite-4.0-H-Micro?
Granite-4.0-H-Micro is text-only, does not provide built-in web search or native image, audio, or video generation, and may be less reliable than larger models for difficult reasoning, complex planning, extensive coding, and broad factual tasks. IBM also does not publish an official hosted token price for this exact model.
Does Granite-4.0-H-Micro support tool calling and multilingual text?
Yes. The instruct variant supports structured tool calling for workflows such as database lookups, calculations, document retrieval, and ticket creation. It supports English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese, although quality may vary by language.
Can Granite-4.0-H-Micro run locally?
Yes. Granite-4.0-H-Micro is distributed as an open-weight model under the Apache 2.0 license and can be deployed with tools such as Transformers, vLLM, SGLang, Docker, and compatible local-inference or quantization systems. Hardware requirements and performance depend on the deployment configuration.
What is IBM Granite-4.0-H-Micro?
IBM Granite-4.0-H-Micro is a 3-billion-parameter open-weight instruct language model designed for locally managed, text-focused applications such as summarization, classification, information extraction, question answering, RAG, coding assistance, and function calling.
How large is the context window of Granite-4.0-H-Micro?
Granite-4.0-H-Micro supports a documented context window of 128,000 tokens, making it suitable for long documents, retrieved passages, code files, and extended conversations. The documentation does not specify a universal maximum output-token limit.


Sources 3
Provider

About IBM watsonx