What Llama Prompt Guard 2 22M is
Llama Prompt Guard 2 22M is a binary text-classification model released by Meta on April 29, 2025. It examines text and assigns it to one of two broad categories: benign or malicious. The malicious category is intended to cover explicit prompt attacks, including attempts to override instructions, jailbreak a model, or inject instructions through content that an application retrieved from an outside source.
In practical terms, the model can sit in front of an LLM-based application. A service might pass a user message, retrieved web page, document excerpt, tool result, or agent input through Prompt Guard 2 22M before including that material in a larger model's context. The classifier's result can then inform an application's filtering, blocking, review, or escalation logic.
The model is not a conversational assistant and does not generate a response to the input. It is a discriminative classifier: its job is to assess text, not to write text, answer questions, call tools, or complete a user task.
Where it fits in the Prompt Guard 2 family
Prompt Guard 2 22M is the smaller member of Meta's Prompt Guard 2 family. Its backbone is DeBERTa-xsmall and it contains approximately 22 million backbone parameters. The family also includes an 86M variant. The 22M model is positioned for lower latency and lower resource use, while the larger sibling is the more appropriate comparison when multilingual performance or additional model capacity is more important.
Unlike Prompt Guard 1, Prompt Guard 2 uses a single benign-or-malicious classification outcome rather than separate injection and jailbreak labels. Applications that require those attack types to be distinguished must therefore add their own analysis or use a different model and decision process.
Core capabilities and input limit
The model's primary capability is recognizing explicit or recognizable prompt-attack patterns. Examples include text that attempts to supersede developer instructions, persuade an assistant to ignore its rules, or hide hostile instructions inside content that an application intends to summarize or analyze.
Its maximum recommended context is 512 tokens. This is an important deployment constraint: a long document or prompt cannot simply be passed to the model as one unrestricted input. Meta recommends splitting longer material into segments and classifying those segments either in parallel or sequentially. The surrounding application must then decide how to combine the segment-level results.
Chunking can make the classifier useful for retrieved documents and web content, but it also creates an application-design issue. An attack may depend on context spread across multiple segments, and separate classification results do not automatically provide a document-level security judgment. Developers should define how they handle borderline scores, multiple flagged chunks, and attacks that rely on interactions between separate pieces of text.
Reported performance and efficiency
Meta reports 19.3 milliseconds of classification latency for 512-token inputs on an A100 GPU. This is a provider-reported measurement rather than a guarantee for every hardware configuration. Actual throughput and latency will vary according to hardware, batching, tokenizer overhead, input length, and the surrounding application.
In Meta's published evaluation, Prompt Guard 2 22M achieved an English area under the curve (AUC) of 0.995. Meta also reports 88.7% English recall at a 1% false-positive rate. Recall describes how many relevant malicious examples are detected, while the false-positive rate describes how often benign examples are incorrectly flagged. These figures should be treated as evaluation results, not as a promise that the same balance will apply to a particular application's traffic.
The main practical advantage of the 22M model is efficiency. A compact classifier can be placed on a local or dedicated inference service and run before an expensive generative model call. The trade-off is that the smaller backbone has weaker multilingual performance than Prompt Guard 2 86M, particularly outside English. Applications serving multiple languages should test the 22M model on representative inputs before relying on it as a primary filter.
Deployment and supported behavior
Meta provides the model as an open-weight download through the Hugging Face repository meta-llama/Llama-Prompt-Guard-2-22M. It can be loaded locally with the Transformers library using a text-classification pipeline or the lower-level AutoTokenizer and AutoModelForSequenceClassification interfaces.
Because it is a local classifier rather than a hosted generative endpoint, the supplied specifications do not list input or output token pricing. Pricing is therefore not applicable to the downloadable model itself, although deployment still has infrastructure, storage, and engineering costs. The model is available under the Llama 4 Community License on Hugging Face, so users should review that license and any applicable usage requirements before deployment.
Prompt Guard 2 22M accepts text and produces a text-classification result. It does not directly accept images, audio, or video, and it does not generate those media types. The supplied specifications also do not identify tool or function calling, streaming, JSON mode, batch API access, or a generative output-token limit. Fine-tuning is listed as supported, which may allow adaptation to application-specific data, but the available research does not provide a fine-tuning procedure or expected improvement.
Main strengths
- Low resource requirements: Approximately 22 million parameters and a DeBERTa-xsmall backbone make it substantially lighter than larger safety or language models.
- Fast screening: Meta reports 19.3 milliseconds for a 512-token classification on an A100 GPU.
- Focused purpose: It is specialized for prompt injections and jailbreak attempts rather than being asked to infer safety from a general-purpose generation task.
- Local deployment: The open-weight model can be integrated into an application without sending every scanned input to a separate hosted moderation service.
- Useful evaluation data: Meta publishes English AUC and recall results, providing more concrete guidance than an unsupported general claim of safety performance.
Limitations and security considerations
Prompt Guard 2 22M is a detection component, not a complete defense. Its predictions can include false positives and false negatives, and attackers may develop adaptive techniques that do not resemble patterns seen during evaluation. A malicious input that passes the classifier should not automatically be treated as safe.
The 512-token limit requires chunking for longer prompts and documents. Chunking may reduce the model's ability to interpret relationships between distant pieces of text, so the application should combine the classifier with access controls, instruction separation, output validation, and other security measures appropriate to its workflow.
Language coverage is another significant limitation. The smaller model is not multilingual by pretraining in the same way as the larger Prompt Guard 2 option, and Meta reports a larger multilingual performance gap for the 22M model. English-focused applications are the clearest fit. Multilingual deployments should measure false positives and missed attacks for each important language rather than assuming the English results transfer directly.
The classifier also provides a broad malicious-versus-benign decision rather than a detailed explanation or a separate label for injection versus jailbreak. Teams that need audit-friendly reasons, attack categories, or policy-specific moderation decisions will need additional logic around it.
When to choose Llama Prompt Guard 2 22M
Choose Llama Prompt Guard 2 22M when the primary requirement is fast, inexpensive screening of text before it reaches an LLM or agent. It is a reasonable fit for user-message filtering, retrieval-augmented generation pipelines, document assistants, web-content ingestion, agent memory, and tool-result inspection. It is especially attractive when local inference, predictable infrastructure costs, or a small deployment footprint matters more than broad language coverage.
The model is also a useful first-stage filter. An application could run this compact classifier on every input and reserve a more capable or expensive security review for suspicious, uncertain, or high-impact cases. That architecture can reduce the number of generative-model calls, although the research does not establish a specific cost reduction.
Choose Prompt Guard 2 86M instead when the application can accept greater resource use and needs the stronger multilingual positioning reported for that sibling. Choose a broader moderation or security system when the requirement includes harmful-content categories, policy enforcement, image or audio analysis, detailed explanations, or protection against sophisticated application-specific attacks. A general-purpose LLM should not be treated as an automatic substitute for this classifier, but Prompt Guard 2 22M should not be treated as a complete substitute for broader controls either.
Practical evaluation guidance
Before production use, test the model with examples from the application's real traffic. Include ordinary user requests, quoted attacks, retrieved documents containing instructions, multilingual inputs, long documents split into chunks, and benign text that resembles a jailbreak. Tune the application's response to a malicious result according to the consequences of blocking legitimate users versus allowing an attack through.
The available catalog assessment rates the model highly for speed and cost efficiency and low for reasoning, while assigning a modest coding score. Those are editorial product-comparison assessments, not benchmarks published by Meta. They reflect the model's role as a small classifier: it is optimized for fast security screening, not for reasoning through complex tasks or writing code.
Overall, Llama Prompt Guard 2 22M is best understood as a lightweight protective layer. Its value comes from placing focused, low-latency classification before a larger model, while its 512-token limit, weaker multilingual performance, binary output, and vulnerability to adaptive attacks define the boundaries of what it can safely do.

