What is Llama Guard 3-1B?
Llama Guard 3-1B is an open-weight language model from Meta that has been fine-tuned for safety classification. Its job is to examine a prompt, a conversation, or a response produced by another language model and determine whether the content should be treated as safe or unsafe. For unsafe content, it can also return the applicable hazard categories.
This makes Llama Guard 3-1B a guardrail component, not a general-purpose assistant. A typical application places it before a primary model to screen incoming requests, after that model to review its responses, or at both points. The model was released on September 25, 2024, alongside Meta’s Llama 3.2 generation, and is part of the PurpleLlama collection of system-level safety tools.
The model is based on Llama 3.2 1B and is intended to reduce the compute and memory cost of adding moderation to an LLM application. Its small size makes local deployment more practical than using a larger safety model or sending every moderation decision to a hosted service.
How the safety classifier works
Llama Guard 3-1B uses a conversational language-model interface, but the expected result is a compact classification rather than a natural conversation. An application supplies text in the format expected by the model, such as a user prompt or a multi-turn exchange. The model then generates a text decision, commonly indicating safe or unsafe. When the result is unsafe, the output can include one or more category labels.
Its default policy follows 13 categories based on the MLCommons hazard taxonomy:
- Violent crimes
- Non-violent crimes
- Sex-related crimes
- Child sexual exploitation
- Defamation
- Specialized advice
- Privacy
- Intellectual property
- Indiscriminate weapons
- Hate
- Suicide and self-harm
- Sexual content
- Elections
These labels give an application more useful information than a single block-or-allow value. For example, a product could send privacy-related cases to a different review process from self-harm cases. The model can also be customized or fine-tuned for an application-specific taxonomy, although the supplied research does not establish a universal performance level for custom policies.
Verified specifications and supported modalities
| Specification | Details |
|---|---|
| Provider | Meta |
| Release date | September 25, 2024 |
| Model size | Approximately 1 billion parameters |
| Reported context length | 131,072 tokens |
| Maximum output | No separate maximum output limit is published in the supplied research |
| Input | Text conversations and prompts |
| Output | Text safety classifications and category labels |
| Languages | English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai |
| Multimodal input | No native image, audio, or video input is documented |
| Multimodal output | No; output is text classification |
| Hosted token pricing | Not applicable as a first-party Meta pay-per-token API model |
The model accepts text and returns text. It does not natively analyze images, listen to audio, process video, or produce media. If an application needs multimodal moderation, another system must first convert those inputs into a form that Llama Guard 3-1B can evaluate, or a separate classifier must be added.
Deployment, context, and cost
Llama Guard 3-1B is downloadable subject to Meta’s access requirements. It can be run locally using the original Llama codebase, Hugging Face Transformers, or compatible serving systems such as vLLM and SGLang. It is therefore different from a conventional hosted model with a published input price and output price.
The supplied model data lists a 131,072-token context length. That is a substantial input capacity for a compact classifier, but applications should still avoid sending irrelevant conversation history. Long inputs increase processing work and may make it harder to apply a consistent policy to the content that matters. The supplied research does not provide a separate maximum number of generated output tokens, so that value should not be assumed from the context length.
Meta provides a pruned and INT4-quantized version for constrained environments. Meta reports that the quantized model’s size was reduced from approximately 2,858 MB to 438 MB. This is a provider-reported storage reduction, not a guarantee of identical accuracy or latency on every device. Quantization can make local deployment more practical, but teams should test classification quality on their own languages, policy categories, and hardware.
Because there is no official Meta input or output token price for this open-weight checkpoint, the effective cost depends on infrastructure. A local deployment may avoid per-request API charges, while still requiring memory, compute, monitoring, and operational maintenance. On-device or edge execution can also reduce the need to transmit sensitive prompts to a third-party service.
Main strengths and trade-offs
The clearest strength of Llama Guard 3-1B is efficiency. Compared with larger safety classifiers, its approximately 1-billion-parameter size and available INT4 version make it a more realistic choice for local, edge, and mobile scenarios. It can provide a dedicated safety layer without using a general-purpose model for every moderation decision.
Its second strength is policy-oriented output. Instead of merely refusing a request, it is designed to identify hazard categories. That structure can support routing, logging, escalation, and different interventions for different risks. The eight-language coverage is also useful for applications operating across the listed languages, although language support should not be treated as proof of equal performance in every category.
The trade-off is capability. The supplied research identifies the larger Llama Guard 3-8B model as generally offering stronger classification performance. Llama Guard 3-1B should therefore be viewed as a lower-cost baseline and fast first-pass filter, not automatically as the best choice for the most sensitive moderation decisions.
Editorial scores in the supplied model data rate its speed and cost favorably, while giving it low reasoning and coding scores. These are comparative editorial assessments, not benchmarks or claims published by Meta. They reflect the model’s intended role: it is optimized for classification, not multi-step reasoning, software development, or open-ended dialogue.
Limitations and safety caveats
Llama Guard 3-1B is not a complete content-governance system. Its decisions can vary with language, category, conversation context, and prompt construction. It may also be less suitable for categories that depend on current factual knowledge, including defamation, intellectual property, and elections. The model card warns that adversarial prompts and prompt-injection techniques can bypass or alter intended behavior.
A classification result should therefore not be treated as an infallible legal, medical, or policy decision. For high-risk uses, developers should combine the model with explicit rules, specialized classifiers where necessary, human review, logging, monitoring, and regular red-team testing. Teams should also verify that Meta’s open-weight license and access terms permit their intended redistribution or commercial deployment.
The model does not provide documented native tool or function calling. It is not a web-search system, embedding model, coding assistant, or general conversational model. The supplied research also does not verify features such as streaming, structured JSON mode, caching, or a batch API. Applications that need those capabilities must implement the surrounding workflow themselves or select a different service.
When to choose Llama Guard 3-1B
Choose Llama Guard 3-1B when the main requirement is inexpensive text safety classification and the team is prepared to operate an open-weight model. It is a strong fit for:
- Screening prompts before they reach a larger language model
- Checking generated responses before displaying them to users
- Local or on-device moderation where privacy, latency, or connectivity matters
- Lightweight safety checks in mobile and edge applications
- Applications that need category labels from the MLCommons hazard taxonomy
- Custom guardrail pipelines that can be tested and fine-tuned for a specific policy
Another option may be more appropriate when moderation quality is more important than memory and operating cost. Meta’s larger Llama Guard 3-8B is the named sibling alternative in the supplied research and generally provides stronger classification performance. A specialized moderation service may also be preferable when a team needs managed infrastructure, continuously updated policies, multimodal moderation, or stronger support processes.
For many applications, the practical design is not to choose between Llama Guard 3-1B and every other safety measure. It can serve as a fast first layer, with rules, a stronger classifier, or human review handling cases where the consequences of a false positive or false negative are significant.
Bottom line
Llama Guard 3-1B is a focused, efficient safety model for text-based LLM guardrails. Its open-weight deployment, eight-language support, 13-category taxonomy, and compact quantized variant make it useful when safety checks need to run close to the application or device. Its limitations are equally important: it is not a general assistant, does not offer native media or tool capabilities, has no official hosted token price, and should not be relied on as the sole control for high-risk moderation. Its best role is a low-cost, testable classification layer within a broader safety system.

