Llama Guard 3

Llama Guard 3-1B

by Meta AI · Current open-weight model; downloadable subject to access approval

Meta Llama Guard 3-1B is a compact open-weight model for classifying prompts and LLM responses as safe or unsafe. It supports eight languages, follows a 13-category MLCommons hazard taxonomy, offers a quantized variant for edge and mobile deployment, and is best used as one layer in a broader safety pipeline.

Text Reasoning Coding
Llama Guard 3-1B is Meta’s small-footprint model for adding safety checks around larger language models. It is designed to classify user inputs and model outputs rather than hold general conversations, write code, use tools, or generate media. Because it can run locally and is available in a quantized form, it is particularly relevant to developers who need inexpensive, low-latency moderation close to the application or device.
Outputs

What Llama Guard 3-1B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

3/10 Reasoning
2/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Llama Guard 3
Model type Other
Context window 131K tokens
Maximum output tokens
Release date 2024-09-25
Status Current open-weight model; downloadable subject to access approval
Knowledge cutoff notes

Meta does not publish a separate knowledge-cutoff date for Llama Guard 3-1B. It is a fine-tuned safety classifier whose performance depends on its underlying Llama 3.2 training data and safety-tuning data.

Model notes

Llama Guard 3-1B is a fine-tuned Llama 3.2 1B model intended to classify LLM inputs and outputs. It generates text labels such as safe or unsafe and can list violated categories from the 13-category MLCommons hazard taxonomy. Meta also provides a pruned and INT4-quantized variant optimized for constrained devices. The model is open-weight rather than a first-party hosted pay-per-token API, so no official Meta input or output token price applies. Meta reports support for English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai. It can be customized or fine-tuned for application-specific taxonomies. The model may be vulnerable to adversarial and prompt-injection attacks, and Meta recommends evaluating it within the complete application safety system.

Model guide

Llama Guard 3-1B: Lightweight Local Safety Classification for LLM Apps

Meta Llama Guard 3-1B is a compact open-weight safety classifier for checking prompts and generated responses. It produces text labels such as safe or unsafe, identifies relevant risks across 13 MLCommons hazard categories, supports eight languages, and includes an INT4 variant for lower-memory edge and mobile deployments.

What is Llama Guard 3-1B?

Llama Guard 3-1B is an open-weight language model from Meta that has been fine-tuned for safety classification. Its job is to examine a prompt, a conversation, or a response produced by another language model and determine whether the content should be treated as safe or unsafe. For unsafe content, it can also return the applicable hazard categories.

This makes Llama Guard 3-1B a guardrail component, not a general-purpose assistant. A typical application places it before a primary model to screen incoming requests, after that model to review its responses, or at both points. The model was released on September 25, 2024, alongside Meta’s Llama 3.2 generation, and is part of the PurpleLlama collection of system-level safety tools.

The model is based on Llama 3.2 1B and is intended to reduce the compute and memory cost of adding moderation to an LLM application. Its small size makes local deployment more practical than using a larger safety model or sending every moderation decision to a hosted service.

How the safety classifier works

Llama Guard 3-1B uses a conversational language-model interface, but the expected result is a compact classification rather than a natural conversation. An application supplies text in the format expected by the model, such as a user prompt or a multi-turn exchange. The model then generates a text decision, commonly indicating safe or unsafe. When the result is unsafe, the output can include one or more category labels.

Its default policy follows 13 categories based on the MLCommons hazard taxonomy:

  • Violent crimes
  • Non-violent crimes
  • Sex-related crimes
  • Child sexual exploitation
  • Defamation
  • Specialized advice
  • Privacy
  • Intellectual property
  • Indiscriminate weapons
  • Hate
  • Suicide and self-harm
  • Sexual content
  • Elections

These labels give an application more useful information than a single block-or-allow value. For example, a product could send privacy-related cases to a different review process from self-harm cases. The model can also be customized or fine-tuned for an application-specific taxonomy, although the supplied research does not establish a universal performance level for custom policies.

Verified specifications and supported modalities

SpecificationDetails
ProviderMeta
Release dateSeptember 25, 2024
Model sizeApproximately 1 billion parameters
Reported context length131,072 tokens
Maximum outputNo separate maximum output limit is published in the supplied research
InputText conversations and prompts
OutputText safety classifications and category labels
LanguagesEnglish, French, German, Hindi, Italian, Portuguese, Spanish, and Thai
Multimodal inputNo native image, audio, or video input is documented
Multimodal outputNo; output is text classification
Hosted token pricingNot applicable as a first-party Meta pay-per-token API model

The model accepts text and returns text. It does not natively analyze images, listen to audio, process video, or produce media. If an application needs multimodal moderation, another system must first convert those inputs into a form that Llama Guard 3-1B can evaluate, or a separate classifier must be added.

Deployment, context, and cost

Llama Guard 3-1B is downloadable subject to Meta’s access requirements. It can be run locally using the original Llama codebase, Hugging Face Transformers, or compatible serving systems such as vLLM and SGLang. It is therefore different from a conventional hosted model with a published input price and output price.

The supplied model data lists a 131,072-token context length. That is a substantial input capacity for a compact classifier, but applications should still avoid sending irrelevant conversation history. Long inputs increase processing work and may make it harder to apply a consistent policy to the content that matters. The supplied research does not provide a separate maximum number of generated output tokens, so that value should not be assumed from the context length.

Meta provides a pruned and INT4-quantized version for constrained environments. Meta reports that the quantized model’s size was reduced from approximately 2,858 MB to 438 MB. This is a provider-reported storage reduction, not a guarantee of identical accuracy or latency on every device. Quantization can make local deployment more practical, but teams should test classification quality on their own languages, policy categories, and hardware.

Because there is no official Meta input or output token price for this open-weight checkpoint, the effective cost depends on infrastructure. A local deployment may avoid per-request API charges, while still requiring memory, compute, monitoring, and operational maintenance. On-device or edge execution can also reduce the need to transmit sensitive prompts to a third-party service.

Main strengths and trade-offs

The clearest strength of Llama Guard 3-1B is efficiency. Compared with larger safety classifiers, its approximately 1-billion-parameter size and available INT4 version make it a more realistic choice for local, edge, and mobile scenarios. It can provide a dedicated safety layer without using a general-purpose model for every moderation decision.

Its second strength is policy-oriented output. Instead of merely refusing a request, it is designed to identify hazard categories. That structure can support routing, logging, escalation, and different interventions for different risks. The eight-language coverage is also useful for applications operating across the listed languages, although language support should not be treated as proof of equal performance in every category.

The trade-off is capability. The supplied research identifies the larger Llama Guard 3-8B model as generally offering stronger classification performance. Llama Guard 3-1B should therefore be viewed as a lower-cost baseline and fast first-pass filter, not automatically as the best choice for the most sensitive moderation decisions.

Editorial scores in the supplied model data rate its speed and cost favorably, while giving it low reasoning and coding scores. These are comparative editorial assessments, not benchmarks or claims published by Meta. They reflect the model’s intended role: it is optimized for classification, not multi-step reasoning, software development, or open-ended dialogue.

Limitations and safety caveats

Llama Guard 3-1B is not a complete content-governance system. Its decisions can vary with language, category, conversation context, and prompt construction. It may also be less suitable for categories that depend on current factual knowledge, including defamation, intellectual property, and elections. The model card warns that adversarial prompts and prompt-injection techniques can bypass or alter intended behavior.

A classification result should therefore not be treated as an infallible legal, medical, or policy decision. For high-risk uses, developers should combine the model with explicit rules, specialized classifiers where necessary, human review, logging, monitoring, and regular red-team testing. Teams should also verify that Meta’s open-weight license and access terms permit their intended redistribution or commercial deployment.

The model does not provide documented native tool or function calling. It is not a web-search system, embedding model, coding assistant, or general conversational model. The supplied research also does not verify features such as streaming, structured JSON mode, caching, or a batch API. Applications that need those capabilities must implement the surrounding workflow themselves or select a different service.

When to choose Llama Guard 3-1B

Choose Llama Guard 3-1B when the main requirement is inexpensive text safety classification and the team is prepared to operate an open-weight model. It is a strong fit for:

  • Screening prompts before they reach a larger language model
  • Checking generated responses before displaying them to users
  • Local or on-device moderation where privacy, latency, or connectivity matters
  • Lightweight safety checks in mobile and edge applications
  • Applications that need category labels from the MLCommons hazard taxonomy
  • Custom guardrail pipelines that can be tested and fine-tuned for a specific policy

Another option may be more appropriate when moderation quality is more important than memory and operating cost. Meta’s larger Llama Guard 3-8B is the named sibling alternative in the supplied research and generally provides stronger classification performance. A specialized moderation service may also be preferable when a team needs managed infrastructure, continuously updated policies, multimodal moderation, or stronger support processes.

For many applications, the practical design is not to choose between Llama Guard 3-1B and every other safety measure. It can serve as a fast first layer, with rules, a stronger classifier, or human review handling cases where the consequences of a false positive or false negative are significant.

Bottom line

Llama Guard 3-1B is a focused, efficient safety model for text-based LLM guardrails. Its open-weight deployment, eight-language support, 13-category taxonomy, and compact quantized variant make it useful when safety checks need to run close to the application or device. Its limitations are equally important: it is not a general assistant, does not offer native media or tool capabilities, has no official hosted token price, and should not be relied on as the sole control for high-risk moderation. Its best role is a low-cost, testable classification layer within a broader safety system.


Answers to Frequently Asked Questions

What are the main limitations of Llama Guard 3-1B?
Llama Guard 3-1B is a text safety classifier rather than a general-purpose assistant or complete content-governance system. Its results can vary by language, category, context, and prompt construction, and adversarial prompts may bypass intended behavior. For high-risk moderation, it should be combined with rules, specialized classifiers, human review, monitoring, and red-team testing.
Does Llama Guard 3-1B support images, audio, or video moderation?
No. Llama Guard 3-1B accepts text conversations and prompts and returns text safety classifications. It has no documented native image, audio, or video input, so multimodal applications need a separate system or must first convert media into text.
Can Llama Guard 3-1B run locally and offline?
Yes. Llama Guard 3-1B can be downloaded and deployed locally using tools such as the Llama codebase, Hugging Face Transformers, vLLM, or SGLang. Meta also provides an INT4-quantized version that reduces the reported model size from approximately 2,858 MB to 438 MB, making local and edge deployment more practical.
What is Llama Guard 3-1B used for?
Llama Guard 3-1B is an open-weight safety classification model from Meta. It evaluates prompts, conversations, and LLM-generated responses to determine whether they are safe or unsafe, and can provide relevant hazard category labels.
Which safety categories does Llama Guard 3-1B support?
Its default policy covers 13 MLCommons-based categories: violent crimes, non-violent crimes, sex-related crimes, child sexual exploitation, defamation, specialized advice, privacy, intellectual property, indiscriminate weapons, hate, suicide and self-harm, sexual content, and elections.


Sources 4
Provider

About Meta AI