What Granite Guardian 3.8B is
Granite Guardian 3.8B is a specialized IBM model for safety and quality evaluation. Instead of serving as the main assistant that writes an answer to a user, it examines prompts, responses, retrieved documents, and other workflow data to identify risks or failures.
The model name can be confusing: despite the “3.8B” label, IBM’s supplied documentation describes the model as an 8-billion-parameter model. It is a fine-tuned version of Granite 3.1 8B Instruct, trained using human annotations, synthetic data, and internal red-teaming data. IBM released it on December 18, 2024, under the Apache 2.0 license.
In practical terms, Granite Guardian can function as a guardrail between a user and a primary language model, or as an evaluator that checks an answer after it has been generated. For example, an application could inspect a user prompt for jailbreak attempts before sending it to a chatbot, then evaluate the chatbot’s response for harmful content or unsupported claims before displaying it.
What risks and quality problems it detects
Granite Guardian 3.8B is intended for configurable oversight across enterprise AI systems. Its documented evaluation areas cover both conventional safety moderation and common failure modes in retrieval-augmented generation, or RAG. RAG systems retrieve information from a supplied knowledge base and provide that material to a language model as context. A response can still be unsafe, irrelevant, or unsupported even when retrieval is enabled, so the model can assess several separate parts of that process.
- Harm-related risks: screening prompts and responses for potentially harmful content.
- Jailbreaks and adversarial prompts: identifying attempts to bypass an application’s safety instructions or restrictions.
- Social bias and profanity: checking for abusive, biased, or profane content.
- Violence, sexual content, and unethical behavior: detecting categories of content that may conflict with an organization’s policies.
- Context relevance: assessing whether retrieved information is relevant to the user’s question.
- Groundedness: checking whether a generated answer is supported by the supplied context.
- Answer relevance: determining whether the response addresses the user’s request.
- Hallucination-related risks: evaluating whether a RAG answer contains claims that are not adequately supported.
IBM’s Granite Guardian documentation also discusses hallucination detection in function-calling workflows for Granite Guardian models in the 3.1 generation. That positioning makes the model relevant to agentic systems, but Granite Guardian 3.8B should not be treated as a tool-calling assistant itself. The supplied model data records tool use as unsupported, so its role is evaluation rather than independently executing functions.
Technical specifications and limits
| Specification | Granite Guardian 3.8B |
|---|---|
| Provider | IBM |
| Model family | Granite Guardian |
| Base model | Granite 3.1 8B Instruct |
| Parameters | 8 billion |
| Architecture | Decoder-only transformer |
| Context limit | 131,072 tokens including input and output |
| Maximum new tokens | 8,192 per request |
| Documented language | English |
| Primary output | Text-based risk assessments and classifications |
| License | Apache 2.0 |
| Release date | December 18, 2024 |
The context limit is the combined allowance for the supplied prompt, any retrieved context, and generated output. It is not a guarantee that every deployment will accept a request of exactly that size: hosting environments can impose their own operational limits. The documented maximum of 8,192 new tokens describes the model’s output allowance per request.
Granite Guardian 3.8B is a text-input, text-output model. The supplied specifications do not document image, audio, or video input, and it does not directly produce images, audio, video, speech, or embeddings. Web search is also not a built-in capability of this exact model.
How it fits in IBM’s catalog
IBM positions Granite Guardian as part of the Granite model family, but it serves a different purpose from a general-purpose instruction model. A general language model is usually selected to write, summarize, classify, or converse. Granite Guardian is selected to inspect the behavior and output of such systems.
IBM’s watsonx.ai foundation-model documentation marks the granite-guardian-3-8b listing as deprecated. That status matters for new deployments: teams considering IBM-hosted inference should check whether the model is still available in their region and environment, and should evaluate newer Granite Guardian releases when they provide the required functionality. The deprecation status does not erase the model’s identity as an open-weight Apache 2.0 checkpoint, but it does make long-term hosted availability less certain.
The model may also be obtained through IBM Granite’s model distribution channels. Self-hosting and IBM-hosted inference are different choices. A watsonx.ai token price covers managed API usage; it does not cover the hardware, memory, serving software, operations, or monitoring required to run the open-weight model yourself.
Pricing and deployment options
IBM’s supplied watsonx.ai developer pricing lists an indicative cost of $0.0002 per 1,000 input tokens and $0.0002 per 1,000 output tokens for Granite Guardian 3.8B. These are usage-based token prices for the documented IBM-hosted API context, not a recurring subscription price.
For an evaluation pipeline, the practical cost depends on how much text is sent to the guard model and how often it is called. A system that checks both the incoming prompt and the outgoing answer may make two evaluation requests for one user interaction. RAG applications can also consume more input tokens because retrieved passages are included in the evaluation context. Although the per-token rates are low in the supplied pricing information, high-volume moderation or repeated agent checks should still be measured using the organization’s actual prompts, contexts, and response lengths.
Self-hosting can provide greater control over data and deployment, but it shifts the cost calculation from token billing to infrastructure and operations. Operators need suitable compute and memory, a compatible serving stack, monitoring, and a process for validating false positives and false negatives. The Apache 2.0 license is permissive, but it does not remove those engineering responsibilities.
Capabilities and trade-offs
Granite Guardian’s main strength is specialization. It is designed to answer questions such as “Is this response supported by the supplied context?” or “Does this prompt attempt to bypass the system’s rules?” That focus makes it more appropriate for a guardrail layer than a general conversational model.
Its 8-billion-parameter size is relatively compact compared with many larger general-purpose language models. The supplied editorial assessment rates its speed at 7 out of 10 and cost at 8 out of 10, reflecting its smaller specialized role rather than a provider-published benchmark. Actual latency and cost will vary with hardware, context size, batching, hosting environment, and request volume.
The same specialization creates important limits. This is not a replacement for a primary model that must carry on an open-ended conversation, write application code, browse the web, interpret images, or process audio and video. Its documented language support is English, so organizations handling other languages should validate performance rather than assume equivalent behavior. Safety classifications are also not universal policy decisions: a company should test the model against its own content rules and measure both false positives and false negatives.
Reasoning and coding should be understood in that context. The supplied editorial reasoning score is 5 out of 10 and the coding score is 2 out of 10; these are subjective database evaluations, not IBM benchmark results. Granite Guardian may analyze whether content is risky or whether an answer is grounded, but it is not intended as a reasoning-first or code-generation model. Likewise, the supplied data marks structured output, streaming, fine-tuning, caching, and batch API support as unverified rather than confirmed capabilities.
Best use cases
Granite Guardian 3.8B is a good fit when an application needs a distinct inspection layer around another model. Suitable examples include:
- Checking user prompts before they reach a customer-service or internal assistant.
- Screening generated responses for harmful content or policy violations.
- Detecting jailbreak attempts in applications exposed to untrusted users.
- Evaluating whether retrieved passages are relevant to a question.
- Checking whether a RAG answer is grounded in the context supplied to the model.
- Adding a safety or quality gate before an answer is shown to a user or passed to another workflow step.
- Monitoring agentic systems where prompts, intermediate results, and final answers require independent evaluation.
A common deployment pattern is to call Granite Guardian before generation for prompt risk detection, after generation for response moderation, or at both stages. In a RAG pipeline, it can also receive the user question, retrieved context, and generated answer so that relevance and groundedness are assessed together.
When to choose Granite Guardian 3.8B
Choose Granite Guardian 3.8B when the main problem is evaluating safety or factual support rather than generating the best possible standalone answer. It is especially relevant for teams that want an open-weight guard model, need Apache 2.0 licensing, or want a separate evaluator for enterprise RAG and moderation workflows.
Another option may be more appropriate when the application needs a general-purpose chat model, strong code generation, multilingual coverage beyond documented English support, native tool execution, web search, or multimodal input. A newer Granite Guardian release may also be preferable for a new IBM deployment if it is available and offers the required capabilities, because the exact Granite Guardian 3.8B watsonx.ai listing is marked deprecated.
Before production use, test the model with representative prompts, retrieved documents, and responses from the target application. Establish acceptable thresholds for each risk category, review borderline classifications, and confirm that the model’s decisions align with organizational policy. Granite Guardian can provide an important evaluation layer, but it should not be treated as an infallible safety authority or as a substitute for application-level controls, human review, and careful system design.

