Granite Guardian

Granite Guardian 3.8B

by IBM watsonx · Deprecated in IBM watsonx.ai documentation; open-weight model and official model materials remain available

Granite Guardian 3.8B is IBM’s specialized 8-billion-parameter safety model for evaluating prompts, responses, retrieved context, and agent workflows. It detects harmful content, jailbreaks, relevance problems, groundedness failures, and hallucination-related risks. The model supports text input and output, offers a 131,072-token context limit, and is available under Apache 2.0, but its watsonx.ai listing is deprecated.

Text Reasoning Coding
Granite Guardian 3.8B is an IBM safety model designed to sit beside a language model and inspect what goes into or comes out of an AI application. It can screen prompts and responses for harmful or adversarial content, and it can evaluate retrieval-augmented generation results for relevance, groundedness, and hallucination-related risks. The model is based on Granite 3.1 8B Instruct, has a documented 131,072-token input-plus-output context limit, and can generate up to 8,192 new tokens per request. It is available under the Apache 2.0 license, although IBM’s watsonx.ai catalog marks this exact model as deprecated.
Outputs

What Granite Guardian 3.8B can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

5/10 Reasoning
2/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Granite Guardian
Model type Other
Context window 131K tokens
Maximum output 8K tokens
Release date 2024-12-18
Status Deprecated in IBM watsonx.ai documentation; open-weight model and official model materials remain available
Knowledge cutoff notes

IBM's available model documentation specifies the release date and training lineage but does not provide a distinct knowledge-cutoff date for this exact model.

Model notes

The canonical IBM watsonx.ai model identifier is granite-guardian-3-8b. IBM describes it as a fine-tuned Granite 3.1 8B Instruct model. The model is designed for risk detection rather than ordinary text generation and supports documented safety and RAG-related evaluation tasks. IBM documentation lists English as the supported natural language. The watsonx.ai foundation-model catalog marks the exact model as deprecated. Pricing applies to IBM-hosted watsonx.ai usage and does not represent self-hosting costs. Editorial scores reflect its specialized guardrail role and should not be compared directly with general-purpose chat models without accounting for task differences.

Cost

Model pricing

Input $0.0002 per 1,000 input tokens in IBM watsonx.ai documentation
Output $0.0002 per 1,000 output tokens in IBM watsonx.ai documentation
Model guide

Granite Guardian 3.8B: IBM’s Open Safety Model for RAG and Jailbreak Detection

Granite Guardian 3.8B is IBM’s specialized safety model for evaluating prompts, generated answers, retrieved context, and agent workflows. It detects harmful content, jailbreak attempts, relevance problems, groundedness failures, and hallucination-related risks rather than acting as a general-purpose chatbot.

What Granite Guardian 3.8B is

Granite Guardian 3.8B is a specialized IBM model for safety and quality evaluation. Instead of serving as the main assistant that writes an answer to a user, it examines prompts, responses, retrieved documents, and other workflow data to identify risks or failures.

The model name can be confusing: despite the “3.8B” label, IBM’s supplied documentation describes the model as an 8-billion-parameter model. It is a fine-tuned version of Granite 3.1 8B Instruct, trained using human annotations, synthetic data, and internal red-teaming data. IBM released it on December 18, 2024, under the Apache 2.0 license.

In practical terms, Granite Guardian can function as a guardrail between a user and a primary language model, or as an evaluator that checks an answer after it has been generated. For example, an application could inspect a user prompt for jailbreak attempts before sending it to a chatbot, then evaluate the chatbot’s response for harmful content or unsupported claims before displaying it.

What risks and quality problems it detects

Granite Guardian 3.8B is intended for configurable oversight across enterprise AI systems. Its documented evaluation areas cover both conventional safety moderation and common failure modes in retrieval-augmented generation, or RAG. RAG systems retrieve information from a supplied knowledge base and provide that material to a language model as context. A response can still be unsafe, irrelevant, or unsupported even when retrieval is enabled, so the model can assess several separate parts of that process.

  • Harm-related risks: screening prompts and responses for potentially harmful content.
  • Jailbreaks and adversarial prompts: identifying attempts to bypass an application’s safety instructions or restrictions.
  • Social bias and profanity: checking for abusive, biased, or profane content.
  • Violence, sexual content, and unethical behavior: detecting categories of content that may conflict with an organization’s policies.
  • Context relevance: assessing whether retrieved information is relevant to the user’s question.
  • Groundedness: checking whether a generated answer is supported by the supplied context.
  • Answer relevance: determining whether the response addresses the user’s request.
  • Hallucination-related risks: evaluating whether a RAG answer contains claims that are not adequately supported.

IBM’s Granite Guardian documentation also discusses hallucination detection in function-calling workflows for Granite Guardian models in the 3.1 generation. That positioning makes the model relevant to agentic systems, but Granite Guardian 3.8B should not be treated as a tool-calling assistant itself. The supplied model data records tool use as unsupported, so its role is evaluation rather than independently executing functions.

Technical specifications and limits

SpecificationGranite Guardian 3.8B
ProviderIBM
Model familyGranite Guardian
Base modelGranite 3.1 8B Instruct
Parameters8 billion
ArchitectureDecoder-only transformer
Context limit131,072 tokens including input and output
Maximum new tokens8,192 per request
Documented languageEnglish
Primary outputText-based risk assessments and classifications
LicenseApache 2.0
Release dateDecember 18, 2024

The context limit is the combined allowance for the supplied prompt, any retrieved context, and generated output. It is not a guarantee that every deployment will accept a request of exactly that size: hosting environments can impose their own operational limits. The documented maximum of 8,192 new tokens describes the model’s output allowance per request.

Granite Guardian 3.8B is a text-input, text-output model. The supplied specifications do not document image, audio, or video input, and it does not directly produce images, audio, video, speech, or embeddings. Web search is also not a built-in capability of this exact model.

How it fits in IBM’s catalog

IBM positions Granite Guardian as part of the Granite model family, but it serves a different purpose from a general-purpose instruction model. A general language model is usually selected to write, summarize, classify, or converse. Granite Guardian is selected to inspect the behavior and output of such systems.

IBM’s watsonx.ai foundation-model documentation marks the granite-guardian-3-8b listing as deprecated. That status matters for new deployments: teams considering IBM-hosted inference should check whether the model is still available in their region and environment, and should evaluate newer Granite Guardian releases when they provide the required functionality. The deprecation status does not erase the model’s identity as an open-weight Apache 2.0 checkpoint, but it does make long-term hosted availability less certain.

The model may also be obtained through IBM Granite’s model distribution channels. Self-hosting and IBM-hosted inference are different choices. A watsonx.ai token price covers managed API usage; it does not cover the hardware, memory, serving software, operations, or monitoring required to run the open-weight model yourself.

Pricing and deployment options

IBM’s supplied watsonx.ai developer pricing lists an indicative cost of $0.0002 per 1,000 input tokens and $0.0002 per 1,000 output tokens for Granite Guardian 3.8B. These are usage-based token prices for the documented IBM-hosted API context, not a recurring subscription price.

For an evaluation pipeline, the practical cost depends on how much text is sent to the guard model and how often it is called. A system that checks both the incoming prompt and the outgoing answer may make two evaluation requests for one user interaction. RAG applications can also consume more input tokens because retrieved passages are included in the evaluation context. Although the per-token rates are low in the supplied pricing information, high-volume moderation or repeated agent checks should still be measured using the organization’s actual prompts, contexts, and response lengths.

Self-hosting can provide greater control over data and deployment, but it shifts the cost calculation from token billing to infrastructure and operations. Operators need suitable compute and memory, a compatible serving stack, monitoring, and a process for validating false positives and false negatives. The Apache 2.0 license is permissive, but it does not remove those engineering responsibilities.

Capabilities and trade-offs

Granite Guardian’s main strength is specialization. It is designed to answer questions such as “Is this response supported by the supplied context?” or “Does this prompt attempt to bypass the system’s rules?” That focus makes it more appropriate for a guardrail layer than a general conversational model.

Its 8-billion-parameter size is relatively compact compared with many larger general-purpose language models. The supplied editorial assessment rates its speed at 7 out of 10 and cost at 8 out of 10, reflecting its smaller specialized role rather than a provider-published benchmark. Actual latency and cost will vary with hardware, context size, batching, hosting environment, and request volume.

The same specialization creates important limits. This is not a replacement for a primary model that must carry on an open-ended conversation, write application code, browse the web, interpret images, or process audio and video. Its documented language support is English, so organizations handling other languages should validate performance rather than assume equivalent behavior. Safety classifications are also not universal policy decisions: a company should test the model against its own content rules and measure both false positives and false negatives.

Reasoning and coding should be understood in that context. The supplied editorial reasoning score is 5 out of 10 and the coding score is 2 out of 10; these are subjective database evaluations, not IBM benchmark results. Granite Guardian may analyze whether content is risky or whether an answer is grounded, but it is not intended as a reasoning-first or code-generation model. Likewise, the supplied data marks structured output, streaming, fine-tuning, caching, and batch API support as unverified rather than confirmed capabilities.

Best use cases

Granite Guardian 3.8B is a good fit when an application needs a distinct inspection layer around another model. Suitable examples include:

  • Checking user prompts before they reach a customer-service or internal assistant.
  • Screening generated responses for harmful content or policy violations.
  • Detecting jailbreak attempts in applications exposed to untrusted users.
  • Evaluating whether retrieved passages are relevant to a question.
  • Checking whether a RAG answer is grounded in the context supplied to the model.
  • Adding a safety or quality gate before an answer is shown to a user or passed to another workflow step.
  • Monitoring agentic systems where prompts, intermediate results, and final answers require independent evaluation.

A common deployment pattern is to call Granite Guardian before generation for prompt risk detection, after generation for response moderation, or at both stages. In a RAG pipeline, it can also receive the user question, retrieved context, and generated answer so that relevance and groundedness are assessed together.

When to choose Granite Guardian 3.8B

Choose Granite Guardian 3.8B when the main problem is evaluating safety or factual support rather than generating the best possible standalone answer. It is especially relevant for teams that want an open-weight guard model, need Apache 2.0 licensing, or want a separate evaluator for enterprise RAG and moderation workflows.

Another option may be more appropriate when the application needs a general-purpose chat model, strong code generation, multilingual coverage beyond documented English support, native tool execution, web search, or multimodal input. A newer Granite Guardian release may also be preferable for a new IBM deployment if it is available and offers the required capabilities, because the exact Granite Guardian 3.8B watsonx.ai listing is marked deprecated.

Before production use, test the model with representative prompts, retrieved documents, and responses from the target application. Establish acceptable thresholds for each risk category, review borderline classifications, and confirm that the model’s decisions align with organizational policy. Granite Guardian can provide an important evaluation layer, but it should not be treated as an infallible safety authority or as a substitute for application-level controls, human review, and careful system design.


Answers to Frequently Asked Questions

How much does Granite Guardian 3.8B cost and is it still available through watsonx.ai?
IBM’s supplied watsonx.ai developer pricing lists an indicative rate of $0.0002 per 1,000 input tokens and $0.0002 per 1,000 output tokens for hosted API usage. The watsonx.ai listing for granite-guardian-3-8b is marked deprecated, so teams should verify regional and environment availability and consider newer Granite Guardian releases for new deployments. Self-hosting has no token subscription fee but requires suitable infrastructure, serving software, monitoring, and operational maintenance.
What are the main technical specifications of Granite Guardian 3.8B?
Granite Guardian 3.8B is an IBM decoder-only transformer based on Granite 3.1 8B Instruct. IBM’s documentation describes it as an 8-billion-parameter, English-language, text-input and text-output model with a 131,072-token combined context limit and a maximum of 8,192 new tokens per request. It was released on December 18, 2024, under the Apache 2.0 license.
Is Granite Guardian 3.8B a general-purpose chatbot or tool-calling model?
No. Granite Guardian 3.8B is an evaluator rather than a primary conversational assistant or tool-calling model. It is intended to assess prompts and responses, not to conduct open-ended conversations, independently execute functions, browse the web, generate code, or process multimodal input.
What is Granite Guardian 3.8B used for?
Granite Guardian 3.8B is an IBM safety and quality evaluation model designed to inspect prompts, generated responses, retrieved documents, and workflow data. It can serve as a guardrail for detecting harmful content, jailbreak attempts, unsupported claims, and other failures in AI and RAG systems.
What risks and RAG quality issues can Granite Guardian 3.8B detect?
The model can evaluate harm-related content, jailbreaks, adversarial prompts, social bias, profanity, violence, sexual content, unethical behavior, context relevance, groundedness, answer relevance, and hallucination-related risks. In RAG workflows, it can assess whether retrieved information is relevant and whether an answer is supported by the supplied context.


Sources 7
Provider

About IBM watsonx