What is Llama Guard 3-8B?
Llama Guard 3-8B is an open-weight text safety classification model provided by Meta through the PurpleLlama project. It is based on the Llama 3.1 8B pretrained model and fine-tuned for a narrower task: evaluating user prompts and language-model responses for potentially policy-violating content.
Unlike a general-purpose language model, Llama Guard 3-8B is intended to work as a safety layer around another model or application. A typical deployment sends a user message, an assistant response, or both to the classifier. The model returns a safe or unsafe result and, when content is unsafe, identifies the relevant hazard category or categories. An application can then block the content, request a safer response, send the case to a reviewer, or apply a more specific internal policy.
The model belongs to the Llama Guard 3 family. The 8B version is the text-focused member described in the supplied model information; Llama Guard 3-11B-Vision is a separate vision variant rather than an image capability of this model.
What the model classifies
Llama Guard 3-8B uses a hazard taxonomy based on MLCommons categories, with an additional category for code-interpreter abuse. Its 14 reported categories cover violent crimes, non-violent crimes, sex-related crimes, child sexual exploitation, defamation, specialized advice, privacy, intellectual property, indiscriminate weapons, hate, suicide and self-harm, sexual content, elections, and code-interpreter abuse.
This makes the model useful for more than simple profanity filtering. For example, an application could use it to screen a request for instructions related to a dangerous activity, check whether a generated answer contains sensitive personal information, or identify a potentially unsafe request directed at a code interpreter. The category output can help an application apply different actions to different risks instead of treating every unsafe result identically.
Meta also reports optimization for safety checks involving search-tool calls and code-interpreter use. This does not mean the model independently operates tools or acts as a tool-calling assistant. Rather, it can help evaluate prompts and responses associated with those tools, including cases where retrieved information or generated code may create additional safety concerns.
Languages, input, and output modalities
The reported supported languages are English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai. This multilingual coverage makes Llama Guard 3-8B more suitable for applications that moderate users or responses across several of these languages than a classifier tested only in English. Developers should still validate performance on their own content, dialects, terminology, and policy definitions.
The model is text-focused. Its verified input and output types are text, and it does not provide native image, audio, or video input or output in this 8B version. An application that needs direct visual moderation should consider a vision-capable safety model or a separate media moderation system rather than assuming that this checkpoint can inspect images.
Technical specifications and deployment
| Specification | Reported detail |
|---|---|
| Provider | Meta |
| Model family | Llama Guard 3 |
| Base model | Llama 3.1 8B |
| Model role | Text safety classifier |
| Context length | 131,072 tokens |
| Supported languages | English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai |
| Weights and hosting | Open-weight and suitable for self-hosting |
| License | Llama 3.2 Community License |
| Quantized checkpoint | INT8 version provided by Meta |
| Maximum output tokens | Not separately published in the supplied model information |
| Hosted API token pricing | Not published in the supplied model information |
The 131,072-token context length is a reported model specification, but it should not be confused with a recommendation to send an entire application transcript on every request. Moderation systems still need practical limits for latency, memory use, and cost. A deployment may choose to classify only the latest user turn, the generated response, selected conversation history, or a policy-relevant excerpt.
Meta also provides an INT8 quantized checkpoint. Quantization reduces the numerical precision used by the model and can reduce memory requirements, which may make local deployment more practical. The trade-off is that teams should test the quantized checkpoint against the full-precision or reference configuration to determine whether classification quality remains acceptable for their use case.
Performance claims and limitations
Meta’s model card reports an English response-classification F1 score of 0.939 and a false-positive rate of 0.040 on an internal evaluation based on the MLCommons hazard taxonomy. F1 combines precision and recall into a single measure, while the false-positive rate describes how often safe examples were incorrectly flagged under the reported evaluation setup. These are provider-reported results, not universal guarantees. They may not predict performance on a particular product’s users, languages, policy thresholds, or adversarial inputs.
The model can make mistakes in areas that require current facts or nuanced judgment. The supplied research specifically highlights defamation, intellectual property, and elections as categories that may require systems with more current factual knowledge. A safety classifier should therefore not be treated as a definitive legal, medical, factual, or political authority.
False positives can disrupt legitimate requests, while false negatives can allow harmful content through. The appropriate threshold depends on the application. A public-facing service for children, for example, may prefer more aggressive blocking, whereas an internal research workflow may route uncertain cases to human review. Teams should measure both error types on representative examples and test adversarial attempts to bypass the policy.
Capabilities and practical trade-offs
Llama Guard 3-8B’s main capability is classification, not open-ended generation. It can produce the safety label and policy-category information needed by a moderation pipeline, but it is not designed to write long answers, solve general reasoning tasks, generate software, or hold a conversation. The supplied catalog assessment rates its reasoning and coding usefulness low because those are not its intended jobs; those ratings are editorial or catalog evaluations, not provider-published benchmark scores.
The same specialization creates a practical cost and speed trade-off. Compared with using a larger general-purpose model to judge safety, an 8B classifier can be a more focused and potentially more economical component when self-hosted, especially with the supplied INT8 checkpoint. However, local deployment still requires suitable infrastructure, monitoring, and maintenance. There is no verified hosted API price in the supplied information, so a direct per-token comparison with commercial moderation APIs cannot be made here.
The model does not provide documented native tool execution, web search, streaming, batch API access, or a separate JSON-mode guarantee in the supplied research. Its outputs can be placed into an application’s own structured pipeline, but developers should not assume that this is equivalent to a provider-guaranteed structured-output feature.
How to use it in a moderation pipeline
- Define the policy. Decide which of the 14 categories matter to the application and what should happen when each category is detected.
- Choose the text to classify. Screen user inputs, model outputs, tool-related requests, or several of these stages depending on the risk model.
- Interpret both the label and category. An unsafe result can trigger blocking, rewriting, escalation, or a request for clarification rather than one universal action.
- Test thresholds and edge cases. Build an evaluation set containing ordinary usage, ambiguous content, multilingual examples, and adversarial attempts.
- Add additional controls. Use application rules, access controls, logging, human review, and domain-specific checks where the model’s taxonomy is insufficient.
For example, a customer-support application might classify the customer’s message before sending it to a general-purpose assistant and then classify the assistant’s proposed response before displaying it. A coding product could add a separate check around code-interpreter requests. In both cases, Llama Guard 3-8B is a guardrail component, not the system’s complete safety architecture.
When to choose Llama Guard 3-8B
Choose Llama Guard 3-8B when you need an open-weight, self-hosted text classifier for screening LLM prompts or responses and want coverage across its eight reported languages. It is particularly relevant when control over model weights, private deployment, customization, or integration with an existing Llama-based stack matters more than turnkey hosted moderation.
- Good fit: self-hosted input and output moderation for open-weight LLM applications.
- Good fit: multilingual text safety checks in the eight supported languages.
- Good fit: baseline detection for unsafe search-tool requests or code-interpreter abuse.
- Good fit: organizations that want to customize policies or fine-tune a starting point for an application-specific taxonomy.
- Less suitable: general-purpose chat, reasoning, coding, or content generation.
- Less suitable: direct image, audio, or video moderation.
- Less suitable: applications that need a fully managed moderation API with published usage pricing.
- Less suitable: decisions requiring authoritative, current factual or legal verification without additional systems.
Another model or service may be more appropriate when the primary requirement is native multimodal moderation, managed infrastructure, real-time factual knowledge, or a provider-supported moderation API with clear service-level guarantees. Within Meta’s safety-model lineup, the separate Llama Guard 3-11B-Vision variant is the more relevant direction for supported visual inputs, while Llama Guard 3-8B remains the text-oriented option.
Bottom line
Llama Guard 3-8B is best understood as a deployable safety filter for language-model systems. Its open weights, broad context window, eight-language coverage, 14-category taxonomy, and INT8 checkpoint make it a practical starting point for teams building their own moderation layer. Its classification results still require application-specific testing, and its lack of native media support, published hosted pricing, and guaranteed factual verification limits where it can be used on its own.

