Granite Guardian

Granite Guardian 4.1 8B

by IBM watsonx · Current and available as an open-weight model

IBM Granite Guardian 4.1 8B is an Apache 2.0 open-weight evaluator for detecting harmful content and jailbreaks, checking RAG relevance and groundedness, identifying function-call hallucinations, and judging custom user-defined criteria. It supports think and no-think modes, an 8,192-token context window, and local deployment through common open-model runtimes.

Text Reasoning Coding
IBM Granite Guardian 4.1 8B is a specialized 8-billion-parameter model designed to act as a guardrail or evaluator inside an AI application rather than as a general-purpose chatbot. It returns tagged judgments about whether content meets a safety, quality, or custom rule, and can run locally under the Apache 2.0 license.
Outputs

What Granite Guardian 4.1 8B can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

6/10 Reasoning
3/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Granite Guardian
Model type Other
Context window 8K tokens
Release date April 29, 2026
Status Current and available as an open-weight model
Knowledge cutoff notes

IBM and the official model card do not provide a directly verifiable knowledge-cutoff date for Granite Guardian 4.1 8B. The model is designed to evaluate supplied prompts, responses, retrieved context, and tool information rather than provide current factual knowledge from its own training data.

Model notes

Granite Guardian 4.1 8B is fine-tuned from ibm-granite/granite-4.1-8b and is intended to be used with its prescribed judging prompt format. It supports think mode with reasoning traces and no-think mode for lower-latency score-only judgments. Outputs are normally emitted using <think> and <score> tags. IBM documents pre-baked criteria for safety, jailbreak, profanity, RAG relevance and groundedness, and function-calling hallucination, while Bring Your Own Criteria supports custom requirements. The model is trained and tested in English and released under Apache 2.0. IBM describes it as a replacement for Granite Guardian 3.3 8B. No official hosted API pricing, maximum output-token limit, batch API, caching feature, or separate JSON-mode capability was verified for this exact downloadable model.

Cost

Model pricing

Input Not applicable; no official hosted token price found for the downloadable model
Output Not applicable; no official hosted token price found for the downloadable model
Model guide

Granite Guardian 4.1 8B: IBM’s Open-Weight AI Safety Evaluator

Granite Guardian 4.1 8B is IBM’s open-weight safety and evaluation model for judging prompts, responses, retrieved context, and function-call content. Fine-tuned from Granite 4.1 8B, it detects risks such as jailbreaks and harmful content, evaluates RAG relevance and groundedness, checks tool-call hallucinations, and supports custom user-defined criteria in think and no-think modes.

What Granite Guardian 4.1 8B does

Granite Guardian 4.1 8B is an open-weight AI safety and evaluation model provided by IBM Research. It is fine-tuned from IBM’s Granite 4.1 8B language model, but its role is narrower: it judges whether prompts, answers, retrieved passages, conversations, and agent-related function calls satisfy defined criteria.

That makes it useful as a checking layer around another language model. For example, an application can send a user prompt to Granite Guardian to look for a jailbreak attempt, inspect a generated answer for harmful content, or compare an answer with retrieved documents to determine whether it is grounded in the supplied context. Its normal result is a yes-or-no style judgment represented with <score> tags, rather than an open-ended conversational response.

IBM describes the 4.1 8B release as a direct replacement for Granite Guardian 3.3 8B. Within IBM’s current model catalog, it is therefore best understood as a specialized guardrail and evaluation model in the Granite family, not as IBM’s general assistant model or a standalone chat product.

Main evaluation capabilities

The model includes pre-trained criteria for several common safety and reliability checks. These criteria allow developers to use it without building a separate classifier for every basic risk category.

  • Safety detection: identifies harmful or unsafe material, including violence, sexual content, profanity, unethical behavior, and other risky content.
  • Jailbreak detection: evaluates whether a prompt attempts to bypass an AI system’s rules or safety controls.
  • RAG evaluation: checks context relevance, answer relevance, and groundedness. Groundedness means whether an answer is supported by the supplied retrieval context rather than invented independently.
  • Function-call evaluation: detects syntactic and semantic hallucinations in agent tool calls. This can help identify calls that are malformed or inconsistent with the information available to the model.
  • Bring Your Own Criteria: lets developers define custom judging rules, such as response formatting, maximum length, domain requirements, or adherence to a particular instruction.
  • Best-of-N ranking: can score multiple candidate answers so an application can select the strongest response for a task with a verifiable criterion.

These functions make Granite Guardian useful in moderation pipelines, model monitoring, response validation, retrieval-augmented generation systems, and agent workflows. The model evaluates supplied material; it is not itself a web-search system, embedding model, or general-purpose knowledge service.

Think mode and no-think mode

Granite Guardian 4.1 8B supports two documented operating modes. In think mode, it produces a reasoning trace inside <think> tags before returning a score inside <score> tags. This can help developers inspect how an evaluation was reached during testing or debugging.

In no-think mode, the model skips the reasoning trace and returns the judgment more directly. This is the practical choice for lower-latency production guardrails when the application mainly needs a score rather than an explanation.

IBM warns that reasoning traces may contain unsafe material and may not always be faithful to the actual decision process. They should not automatically be treated as a reliable explanation or exposed to end users without filtering. In production, the score should be treated as the primary output, with validation around both the score format and any optional trace.

Technical specifications and supported modalities

The verified configuration documented for this model has 8 billion parameters and an 8,192-token context window. The context window is the maximum amount of input context the model can process in one request, including the judging instructions and the material being evaluated. Developers should account for that limit when passing long retrieved documents, conversations, or multiple candidate responses.

SpecificationVerified detail
ProviderIBM
Model familyGranite Guardian
Parameters8 billion
Context window8,192 tokens
InputText
OutputText-based tagged judgments
LicenseApache 2.0
Language focusEnglish
Maximum output tokensNot verified

The model supports text input and text output only. It does not provide image, audio, video, music, speech, or embedding output. It also does not generate images or audio. Although it can evaluate information describing function calls, the supplied research does not verify built-in tool execution or a tool-use API; the application remains responsible for executing and validating any external function.

Deployment, pricing, and availability

Granite Guardian 4.1 8B is available as downloadable weights through IBM’s Granite collection on Hugging Face. The documented license is Apache 2.0, and the model can be run locally or on infrastructure managed by the operator. Supported deployment options listed in the research include Transformers, vLLM, SGLang, Ollama, Docker, and compatible quantized runtimes.

There is no verified official hosted token price for this exact downloadable model. Consequently, it does not have a meaningful standard monthly or per-token price in the supplied research. The practical cost depends on the hardware, runtime, quantization choice, throughput requirements, and whether the operator uses local or rented infrastructure. Running an 8-billion-parameter model can be more economical and controllable than sending every evaluation to a hosted general-purpose model, but the actual result depends on deployment scale.

IBM lists the release date as April 29, 2026. The model is described as current and available as an open-weight release, although model availability and runtime compatibility can change as hosting platforms and Granite releases are updated.

Strengths and trade-offs

The most important strength is specialization. A model designed specifically to judge safety, groundedness, relevance, and custom criteria can be easier to integrate into a predictable checking pipeline than a general conversational model prompted to act as a safety reviewer. Its open weights also give teams more control over where inference runs and how prompts and evaluation data are handled.

The 8B size is a compromise between evaluator capability and operating cost. It is smaller and potentially faster to run than a large general-purpose judge, making it suitable for repeated checks on incoming prompts and outgoing responses. The no-think mode can further reduce latency when an application needs a binary decision. These are practical trade-offs rather than provider-published performance guarantees; the supplied research does not provide a universal latency or accuracy benchmark.

There are also important limitations. The model is intended to be used with IBM’s prescribed judging prompt format. Deviating from that format, especially under adversarial prompting, can produce unexpected or unsafe behavior. The 8,192-token context limit may require applications to select or summarize long documents before evaluation. No maximum output-token limit, caching feature, batch API, or separate JSON mode was verified for this exact downloadable model.

Its English-focused training and testing also make it a less certain choice for multilingual safety systems. Teams operating in other languages should validate performance for their specific content instead of assuming that the documented criteria transfer equally well.

Reasoning, coding, and tool-related support

Granite Guardian’s reasoning capability is evaluative rather than conversational. Think mode can produce a reasoning trace for inspection, while no-think mode provides a more direct score. This does not make the model a general reasoning assistant, and the research does not establish that it should be used for complex planning or open-ended problem solving.

It can judge custom rules applied to code or technical responses, but it is not positioned as a code-generation model. The supplied evaluation rates its coding capability as limited, an editorial assessment rather than an IBM-published benchmark or specification. Similarly, its function-call feature concerns detecting hallucinated calls; it should not be confused with verified native tool execution.

When to choose Granite Guardian 4.1 8B

Choose this model when the main task is to evaluate another AI system’s behavior and you want an open-weight component that can run under your control. Strong use cases include:

  • blocking or flagging jailbreak attempts and harmful prompts;
  • checking generated answers for safety before delivery;
  • testing whether RAG answers are relevant to and grounded in retrieved documents;
  • validating whether agent function calls are syntactically and semantically plausible;
  • enforcing custom response rules such as required structure, length, or domain-specific instructions;
  • ranking several candidate responses in a best-of-N workflow; and
  • running repeated evaluations in local, private, or cost-sensitive infrastructure.

Another option may be more appropriate when you need an open-ended assistant, reliable multilingual coverage, image or audio processing, web search, embeddings, or native tool execution. A larger judge model may also be preferable when the evaluation is unusually subtle, although that can increase latency and infrastructure or API costs. Conversely, a smaller classifier or rules-based check may be more efficient for a narrow, deterministic policy. Granite Guardian is most compelling when one model needs to cover several safety and quality criteria while remaining deployable as an open-weight text evaluator.

Bottom line

Granite Guardian 4.1 8B is a focused safety and quality-control model, not a replacement for a general-purpose language model. Its combination of pre-built risk checks, RAG evaluation, function-call hallucination detection, custom criteria, think and no-think modes, and Apache 2.0 weights makes it a practical candidate for guardrails around chatbots, retrieval systems, and AI agents. The main implementation responsibilities remain prompt-format compliance, output validation, language testing, context management, and choosing infrastructure that meets the required latency and cost.


Answers to Frequently Asked Questions

How can Granite Guardian 4.1 8B be deployed and what does it cost?
The model is available as downloadable weights through IBM’s Granite collection on Hugging Face and can run locally or on operator-managed infrastructure using options such as Transformers, vLLM, SGLang, Ollama, Docker, and compatible quantized runtimes. No official hosted token price was verified; costs depend on hardware, runtime, quantization, throughput, and infrastructure.
What are the technical specifications and license of Granite Guardian 4.1 8B?
Granite Guardian 4.1 8B has 8 billion parameters, an 8,192-token context window, text-only input and output, an English language focus, and an Apache 2.0 license. It returns text-based tagged judgments rather than images, audio, embeddings, or other non-text outputs.
What are think mode and no-think mode in Granite Guardian 4.1 8B?
In think mode, the model produces a reasoning trace inside tags before returning a judgment inside tags. In no-think mode, it skips the reasoning trace and returns the judgment more directly, which can reduce latency in production guardrails. IBM cautions that reasoning traces may be unsafe or unfaithful explanations.
What is Granite Guardian 4.1 8B designed to do?
Granite Guardian 4.1 8B is an open-weight AI safety and evaluation model from IBM Research. It evaluates prompts, generated answers, retrieved passages, conversations, and agent function calls against safety, reliability, relevance, groundedness, or custom criteria.
What evaluation capabilities does Granite Guardian 4.1 8B support?
It supports safety and harmful-content detection, jailbreak detection, RAG evaluation for context relevance, answer relevance, and groundedness, function-call evaluation, custom criteria through Bring Your Own Criteria, and best-of-N candidate ranking.


Sources 3
Provider

About IBM watsonx