What is Llama Guard 4?
Llama Guard 4 is a 12-billion-parameter dense safety classifier developed by Meta. It is designed to moderate both sides of an AI interaction: the user’s input and the response produced by a generative model. This makes it suitable for moderation pipelines in chatbots, image-aware assistants, content-generation tools, and other applications where unsafe prompts or outputs need to be detected.
The model accepts text-only prompts and mixed text-and-image inputs, including prompts containing multiple images. Its output is textual rather than visual or conversational. In practical terms, an application sends content to Llama Guard 4 and receives a safety decision, along with relevant hazard labels when the content is classified as unsafe.
Meta describes Llama Guard 4 as being derived from the shared dense expert of Llama 4 Scout. Meta removed the routed experts and router layers from the original mixture-of-experts architecture, then fine-tuned the resulting model for content-safety classification. Llama Guard 4 therefore belongs to the Llama Guard safety-model line, not to the general-purpose assistant category.
What can Llama Guard 4 classify?
Llama Guard 4 follows the MLCommons hazard taxonomy and adds a text-only category for Code Interpreter Abuse. Its 14 categories cover a broad range of risks:
- Violent crime
- Non-violent crime
- Sex-related crime
- Child sexual exploitation
- Defamation
- Specialized advice
- Privacy
- Intellectual property
- Indiscriminate weapons
- Hate
- Suicide and self-harm
- Sexual content
- Elections
- Code Interpreter Abuse
The categories can be used to distinguish different policy concerns rather than treating every unsafe result as the same kind of violation. Code Interpreter Abuse is intended for text-only tool-call scenarios in which a request or proposed action could misuse an execution environment.
Inputs, outputs and context length
The verified context length is 8,192 tokens. The model supports text input and image input, including multiple images in a single prompt. The available research does not provide a separate maximum image count or a more detailed image-resolution limit, so those constraints should be checked in the model’s implementation and serving environment.
Llama Guard 4 produces text safety labels. It does not generate images, audio, video, embeddings, or other non-text outputs. Its output is intended to be consumed by application logic, such as a filter, routing rule, review queue, or escalation workflow. The supplied specifications do not identify a guaranteed maximum output-token limit.
| Specification | Verified information |
|---|---|
| Provider | Meta |
| Model size | 12 billion parameters |
| Model type | Dense safety classifier |
| Text input | Yes |
| Image input | Yes, including multiple images |
| Audio, video and image output | No |
| Text output | Yes; safety labels and hazard information |
| Context length | 8,192 tokens |
| Tool or function use | Not supported as a model capability |
| Hosted per-token price | No official first-party price verified for the exact model |
How to use it in a moderation pipeline
Llama Guard 4 can be placed at one or both ends of a generative AI workflow. For input moderation, an application sends the user’s prompt to the classifier before forwarding it to a main language or vision-language model. If the result is unsafe, the application can refuse the request, provide a safer response, or send it for human review.
For output moderation, the application sends the generated answer to Llama Guard 4 before displaying it. This is useful when the original request appeared harmless but the generation nevertheless produced content that violates the application’s policy. Running checks on both input and output provides a defense-in-depth design, although it does not guarantee that every unsafe case will be detected.
A typical workflow might look like this:
- Receive a user prompt, optionally containing images.
- Run the prompt through Llama Guard 4.
- Block, revise, route, or allow the request according to the returned label and hazard category.
- Send permitted content to the generative model.
- Run the generated response through Llama Guard 4 before returning it.
- Log uncertain or high-impact cases for monitoring and human review.
Because the model returns classification results rather than enforcing an application’s complete policy, developers still need to define how each category is handled. A production system may also need rate limits, audit logging, appeal processes, regional policy rules, and additional specialist checks.
Deployment and access
Meta distributes Llama Guard 4 as open weights through its Llama ecosystem, including the official Meta Llama organization on Hugging Face. The model is gated: users must accept the applicable Llama 4 Community License and use policy before downloading the weights.
Meta and Hugging Face document compatibility with Transformers, vLLM, SGLang, Docker-based tooling, and other inference environments. The release documentation states that the model can run on a single GPU with approximately 24 GB of VRAM. This is a provider or release-documentation claim rather than a universal hardware guarantee. Actual memory requirements can change with numerical precision, batching, image count, sequence length, and serving configuration.
There is no verified official first-party per-token hosted API price for the exact Llama Guard 4 model. The economic trade-off is therefore different from that of a conventional hosted model with a published input and output rate: organizations may avoid per-token provider charges by running the weights themselves, but they must account for GPU capacity, deployment, monitoring, maintenance, and engineering costs.
Reasoning, coding and tool-use trade-offs
Llama Guard 4 is not intended to perform open-ended reasoning, coding assistance, factual research, or autonomous task execution. It can classify requests involving these subjects, and its Code Interpreter Abuse category can help identify certain risky text-only tool-call scenarios, but the model itself is not a code-execution engine or tool-using assistant.
Its narrow purpose is an advantage when an application needs a dedicated moderation layer rather than another general-purpose model. It is also a limitation: a classification label should not be treated as a complete explanation of intent, factual truth, legal status, or user risk. Applications that need nuanced policy interpretation may need additional models, deterministic rules, specialist classifiers, or human review.
Main strengths and limitations
Strengths
- Multimodal moderation: It can evaluate text, text with one image, and text with multiple images.
- Input and output coverage: It can screen both user prompts and generated responses.
- Broad hazard taxonomy: Its 14 categories cover safety, privacy, intellectual-property, election-related, and tool-abuse concerns.
- Open-weight deployment: Organizations can download and operate the model in supported environments instead of relying exclusively on a hosted moderation endpoint.
- Dedicated role: Its output is designed for straightforward integration into filtering and routing logic.
Limitations
- Not a general assistant: It should not replace a conversational model, research system, coding model, or policy engine.
- Classification is not perfect judgment: Results can be affected by training-data limitations, multilingual coverage, policy interpretation, common-sense reasoning, and adversarial prompting.
- Some categories need current information: Elections, defamation, and intellectual property may require up-to-date facts or more complex moderation systems than a static classifier can provide.
- Adversarial vulnerability: Prompt injection and other attacks can reduce reliability, particularly when users deliberately attempt to evade moderation.
- Image-distribution limits: Performance may vary when prompts contain substantially more images than those represented in evaluation data.
- Operational cost remains: Open weights remove dependence on a specific hosted price but do not make GPU hosting, integration, or review operations free.
When to choose Llama Guard 4
Choose Llama Guard 4 when you need a dedicated safety checkpoint for a text or vision-language application and want the option to deploy open weights. It is particularly relevant for systems that need to moderate image-aware prompts, inspect generated responses, or apply the same safety layer before and after generation.
It can be a good fit when self-hosting, data-control requirements, or integration flexibility matter more than using a managed moderation API. Its single-GPU deployment claim may also make it practical for smaller deployments, although the real hardware requirement must be tested against the application’s latency and throughput targets.
Another option may be more appropriate when the primary need is general-purpose conversation, advanced reasoning, coding, web research, or reliable tool execution. A hosted moderation service may be preferable when a team does not want to operate GPUs or maintain model infrastructure. High-stakes decisions should not rely on Llama Guard 4 alone; they require application-specific validation, monitoring, clear policies, and human escalation.
Bottom line
Llama Guard 4 is best understood as a moderation component rather than an assistant. Its main distinction is the combination of open-weight deployment, text-and-image input support, multiple-image handling, and a 14-category safety taxonomy. Those capabilities make it useful around generative models, but its labels are only one part of a responsible safety system. Developers should validate it against their own content, languages, policies, attack patterns, and operating conditions before relying on it in production.

