Llama Prompt Guard 2

Llama Prompt Guard 2 22M

by Meta AI · Current open-weight safety classifier

Meta's Llama Prompt Guard 2 22M is a compact open-weight classifier that labels text as benign or malicious to help detect prompt injections and jailbreak attempts. Its 512-token limit, DeBERTa-xsmall backbone, local Transformers deployment, and reported 19.3 ms A100 latency make it suitable for fast screening, while weaker multilingual performance and limited attack context constrain its use.

Text Reasoning Coding
Llama Prompt Guard 2 22M is an open-weight safety classifier from Meta that labels text as benign or malicious. Its purpose is to detect prompt injections and jailbreak attempts before untrusted content reaches an LLM or agent. With approximately 22 million parameters, a 512-token context limit, and reported low-latency inference, it is designed for applications that need inexpensive, local screening rather than text generation.
Outputs

What Llama Prompt Guard 2 22M can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

1/10 Reasoning
2/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Llama Prompt Guard 2
Model type Other
Context window 512 tokens
Maximum output tokens
Release date 2025-04-29
Status Current open-weight safety classifier
Knowledge cutoff notes

This is a discriminative safety classifier rather than a general-purpose knowledge model. Meta does not publish a conventional knowledge-cutoff date for this exact model.

Model notes

The exact Hugging Face model ID is meta-llama/Llama-Prompt-Guard-2-22M. It is a binary text classifier that labels inputs as benign or malicious; Prompt Guard 2 does not retain Prompt Guard 1's separate injection and jailbreak labels. The model is based on DeBERTa-xsmall and has approximately 22 million backbone parameters. Meta reports 19.3 ms latency per 512-token classification on an A100 GPU, English AUC of 0.995, and recall of 88.7% at a 1% false-positive rate. Meta recommends splitting inputs longer than 512 tokens into chunks. The model is available under the Llama 4 Community License on Hugging Face. Pricing is not applicable to the downloadable model itself.

Model guide

Llama Prompt Guard 2 22M: Fast, Lightweight Detection for Prompt Attacks

Llama Prompt Guard 2 22M is Meta's compact, open-weight text classifier for identifying malicious prompt injections and jailbreak attempts. Built on DeBERTa-xsmall, it is the smaller and faster Prompt Guard 2 model, supporting 512-token inputs with lower compute requirements than the 86M variant. It is intended as a screening layer for LLM applications, agents, retrieved content, tool results, and other untrusted text rather than as a generative model or complete security solution.

What Llama Prompt Guard 2 22M is

Llama Prompt Guard 2 22M is a binary text-classification model released by Meta on April 29, 2025. It examines text and assigns it to one of two broad categories: benign or malicious. The malicious category is intended to cover explicit prompt attacks, including attempts to override instructions, jailbreak a model, or inject instructions through content that an application retrieved from an outside source.

In practical terms, the model can sit in front of an LLM-based application. A service might pass a user message, retrieved web page, document excerpt, tool result, or agent input through Prompt Guard 2 22M before including that material in a larger model's context. The classifier's result can then inform an application's filtering, blocking, review, or escalation logic.

The model is not a conversational assistant and does not generate a response to the input. It is a discriminative classifier: its job is to assess text, not to write text, answer questions, call tools, or complete a user task.

Where it fits in the Prompt Guard 2 family

Prompt Guard 2 22M is the smaller member of Meta's Prompt Guard 2 family. Its backbone is DeBERTa-xsmall and it contains approximately 22 million backbone parameters. The family also includes an 86M variant. The 22M model is positioned for lower latency and lower resource use, while the larger sibling is the more appropriate comparison when multilingual performance or additional model capacity is more important.

Unlike Prompt Guard 1, Prompt Guard 2 uses a single benign-or-malicious classification outcome rather than separate injection and jailbreak labels. Applications that require those attack types to be distinguished must therefore add their own analysis or use a different model and decision process.

Core capabilities and input limit

The model's primary capability is recognizing explicit or recognizable prompt-attack patterns. Examples include text that attempts to supersede developer instructions, persuade an assistant to ignore its rules, or hide hostile instructions inside content that an application intends to summarize or analyze.

Its maximum recommended context is 512 tokens. This is an important deployment constraint: a long document or prompt cannot simply be passed to the model as one unrestricted input. Meta recommends splitting longer material into segments and classifying those segments either in parallel or sequentially. The surrounding application must then decide how to combine the segment-level results.

Chunking can make the classifier useful for retrieved documents and web content, but it also creates an application-design issue. An attack may depend on context spread across multiple segments, and separate classification results do not automatically provide a document-level security judgment. Developers should define how they handle borderline scores, multiple flagged chunks, and attacks that rely on interactions between separate pieces of text.

Reported performance and efficiency

Meta reports 19.3 milliseconds of classification latency for 512-token inputs on an A100 GPU. This is a provider-reported measurement rather than a guarantee for every hardware configuration. Actual throughput and latency will vary according to hardware, batching, tokenizer overhead, input length, and the surrounding application.

In Meta's published evaluation, Prompt Guard 2 22M achieved an English area under the curve (AUC) of 0.995. Meta also reports 88.7% English recall at a 1% false-positive rate. Recall describes how many relevant malicious examples are detected, while the false-positive rate describes how often benign examples are incorrectly flagged. These figures should be treated as evaluation results, not as a promise that the same balance will apply to a particular application's traffic.

The main practical advantage of the 22M model is efficiency. A compact classifier can be placed on a local or dedicated inference service and run before an expensive generative model call. The trade-off is that the smaller backbone has weaker multilingual performance than Prompt Guard 2 86M, particularly outside English. Applications serving multiple languages should test the 22M model on representative inputs before relying on it as a primary filter.

Deployment and supported behavior

Meta provides the model as an open-weight download through the Hugging Face repository meta-llama/Llama-Prompt-Guard-2-22M. It can be loaded locally with the Transformers library using a text-classification pipeline or the lower-level AutoTokenizer and AutoModelForSequenceClassification interfaces.

Because it is a local classifier rather than a hosted generative endpoint, the supplied specifications do not list input or output token pricing. Pricing is therefore not applicable to the downloadable model itself, although deployment still has infrastructure, storage, and engineering costs. The model is available under the Llama 4 Community License on Hugging Face, so users should review that license and any applicable usage requirements before deployment.

Prompt Guard 2 22M accepts text and produces a text-classification result. It does not directly accept images, audio, or video, and it does not generate those media types. The supplied specifications also do not identify tool or function calling, streaming, JSON mode, batch API access, or a generative output-token limit. Fine-tuning is listed as supported, which may allow adaptation to application-specific data, but the available research does not provide a fine-tuning procedure or expected improvement.

Main strengths

  • Low resource requirements: Approximately 22 million parameters and a DeBERTa-xsmall backbone make it substantially lighter than larger safety or language models.
  • Fast screening: Meta reports 19.3 milliseconds for a 512-token classification on an A100 GPU.
  • Focused purpose: It is specialized for prompt injections and jailbreak attempts rather than being asked to infer safety from a general-purpose generation task.
  • Local deployment: The open-weight model can be integrated into an application without sending every scanned input to a separate hosted moderation service.
  • Useful evaluation data: Meta publishes English AUC and recall results, providing more concrete guidance than an unsupported general claim of safety performance.

Limitations and security considerations

Prompt Guard 2 22M is a detection component, not a complete defense. Its predictions can include false positives and false negatives, and attackers may develop adaptive techniques that do not resemble patterns seen during evaluation. A malicious input that passes the classifier should not automatically be treated as safe.

The 512-token limit requires chunking for longer prompts and documents. Chunking may reduce the model's ability to interpret relationships between distant pieces of text, so the application should combine the classifier with access controls, instruction separation, output validation, and other security measures appropriate to its workflow.

Language coverage is another significant limitation. The smaller model is not multilingual by pretraining in the same way as the larger Prompt Guard 2 option, and Meta reports a larger multilingual performance gap for the 22M model. English-focused applications are the clearest fit. Multilingual deployments should measure false positives and missed attacks for each important language rather than assuming the English results transfer directly.

The classifier also provides a broad malicious-versus-benign decision rather than a detailed explanation or a separate label for injection versus jailbreak. Teams that need audit-friendly reasons, attack categories, or policy-specific moderation decisions will need additional logic around it.

When to choose Llama Prompt Guard 2 22M

Choose Llama Prompt Guard 2 22M when the primary requirement is fast, inexpensive screening of text before it reaches an LLM or agent. It is a reasonable fit for user-message filtering, retrieval-augmented generation pipelines, document assistants, web-content ingestion, agent memory, and tool-result inspection. It is especially attractive when local inference, predictable infrastructure costs, or a small deployment footprint matters more than broad language coverage.

The model is also a useful first-stage filter. An application could run this compact classifier on every input and reserve a more capable or expensive security review for suspicious, uncertain, or high-impact cases. That architecture can reduce the number of generative-model calls, although the research does not establish a specific cost reduction.

Choose Prompt Guard 2 86M instead when the application can accept greater resource use and needs the stronger multilingual positioning reported for that sibling. Choose a broader moderation or security system when the requirement includes harmful-content categories, policy enforcement, image or audio analysis, detailed explanations, or protection against sophisticated application-specific attacks. A general-purpose LLM should not be treated as an automatic substitute for this classifier, but Prompt Guard 2 22M should not be treated as a complete substitute for broader controls either.

Practical evaluation guidance

Before production use, test the model with examples from the application's real traffic. Include ordinary user requests, quoted attacks, retrieved documents containing instructions, multilingual inputs, long documents split into chunks, and benign text that resembles a jailbreak. Tune the application's response to a malicious result according to the consequences of blocking legitimate users versus allowing an attack through.

The available catalog assessment rates the model highly for speed and cost efficiency and low for reasoning, while assigning a modest coding score. Those are editorial product-comparison assessments, not benchmarks published by Meta. They reflect the model's role as a small classifier: it is optimized for fast security screening, not for reasoning through complex tasks or writing code.

Overall, Llama Prompt Guard 2 22M is best understood as a lightweight protective layer. Its value comes from placing focused, low-latency classification before a larger model, while its 512-token limit, weaker multilingual performance, binary output, and vulnerability to adaptive attacks define the boundaries of what it can safely do.


Answers to Frequently Asked Questions

How can Llama Prompt Guard 2 22M be deployed?
The open-weight model is available from the Hugging Face repository meta-llama/Llama-Prompt-Guard-2-22M and can be run locally with the Transformers library using a text-classification pipeline, AutoTokenizer, or AutoModelForSequenceClassification. Longer inputs should be split into segments, and the application must define how to combine the results.
How does Llama Prompt Guard 2 22M compare with Prompt Guard 2 86M?
Llama Prompt Guard 2 22M is the smaller and more resource-efficient model, using a DeBERTa-xsmall backbone with approximately 22 million backbone parameters. Prompt Guard 2 86M requires more resources but is the better option when stronger multilingual performance or additional model capacity is important.
How fast and accurate is Llama Prompt Guard 2 22M?
Meta reports 19.3 milliseconds of classification latency for 512-token inputs on an A100 GPU. In Meta's published evaluation, it achieved an English AUC of 0.995 and 88.7% English recall at a 1% false-positive rate. Actual results depend on hardware, input length, batching, and application design.
What are the main limitations of Llama Prompt Guard 2 22M?
The model has a recommended maximum context of 512 tokens, produces only a broad benign-or-malicious result, and may generate false positives or false negatives. Its multilingual performance is weaker than Prompt Guard 2 86M, and it should be used as one security layer rather than a complete defense against prompt attacks.
What is Llama Prompt Guard 2 22M used for?
Llama Prompt Guard 2 22M is a binary text-classification model that detects likely prompt attacks, including instruction overrides, jailbreak attempts, and malicious instructions embedded in retrieved content. It can screen user messages, documents, web pages, tool results, and agent inputs before they reach an LLM.


Sources 3
Provider

About Meta AI