What Llama Prompt Guard 2 86M is designed to do
Llama Prompt Guard 2 86M is an open-weight text-classification model from Meta's Llama organization. Its specific job is to identify malicious prompts that could manipulate an LLM-powered application. It is intended to sit around a larger model as a screening layer, not to generate answers or replace the model performing the main task.
The classifier focuses on two closely related attack categories. A prompt injection attempts to make a system ignore or override its existing instructions, often by placing hostile directions inside user input, documents, web pages, emails, retrieved passages, or tool results. A jailbreak is an attempt to bypass a model's built-in safety behavior through specially constructed instructions. Prompt Guard 2 represents both categories with a malicious classification rather than returning separate injection and jailbreak labels.
In practical terms, an application can scan untrusted text before passing it to an LLM, agent, retrieval-augmented generation pipeline, or tool. The application then decides what to do with the result: allow the content, block it, request review, isolate it, or apply a stricter processing policy.
Position in Meta's model lineup
Prompt Guard 2 86M sits in Meta's Llama model catalog as a focused safety component. It is not a general-purpose Llama language model and does not produce conversational, reasoning, coding, image, audio, or video output. Its value comes from specialization: it performs a narrow security classification task that can be added to systems using other models.
The model is published through Meta's official Llama organization on Hugging Face. The canonical identifier is meta-llama/Llama-Prompt-Guard-2-86M. The model is listed as a current open-weight model, but download access is gated and requires acceptance of the license terms displayed on Hugging Face.
Capabilities and architecture
The model contains approximately 86 million parameters and uses a multilingual DeBERTa-v3 backbone. Its classification output is a binary decision: BENIGN or MALICIOUS. Model logits are also available when using the Transformers implementation, allowing an application to apply its own decision threshold instead of relying only on a default class choice.
According to the supplied model information, Prompt Guard 2 uses expanded training data, a refined training objective intended to reduce false positives, and tokenization improvements designed to make attacks involving unusual whitespace or fragmented tokens more difficult. These are provider or model-card claims and should not be treated as a guarantee that every such attack will be detected.
The model card reports evaluation in English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai. This makes it more suitable for multilingual screening than a detector evaluated only on English traffic, although results can still vary by language, application, and attack style.
Technical specifications and supported data types
| Specification | Details |
|---|---|
| Provider | Meta |
| Model family | Llama Prompt Guard 2 |
| Model type | Multilingual text classifier for safety detection |
| Parameters | 86 million |
| Context length | 512 tokens |
| Input | Text only |
| Output | Benign or malicious classification, with logits available through Transformers |
| Deployment | Local or self-hosted inference through Hugging Face Transformers |
| Evaluated languages | English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai |
| Release date | April 29, 2025 |
Prompt Guard 2 does not accept images, audio, or video, and it does not generate those modalities. It also does not provide general text generation. Its text output is a classification result, so it should be integrated as a detector rather than called as a chat model.
The 512-token limit and long documents
The model has a 512-token context window. A token is a unit used by the model to process text and may represent a word, part of a word, punctuation, or whitespace. The limit is short compared with the length of many documents, tool outputs, and web pages.
Inputs longer than 512 tokens should be divided into segments and scanned separately. A security layer can then combine the segment-level results according to its policy. For example, it might block a document if any segment is classified as malicious, assign a risk score based on the highest segment score, or send uncertain cases to human review.
Chunking improves compatibility with the model but creates an application-design responsibility. Attack instructions may be split across segments, and independently classifying each segment can lose some surrounding context. Applications should therefore test chunk size, overlap, aggregation rules, and thresholds using realistic benign and malicious examples.
Reported performance and threshold selection
The official model information reports an English AUC of 0.998 and a multilingual AUC of 0.995 on its direct jailbreak-detection evaluation. AUC, or area under the receiver operating characteristic curve, summarizes how well a classifier separates two classes across different thresholds. The model card also reports 97.5% English recall at a 1% false-positive rate in that evaluation.
These figures describe the reported evaluation setup, not guaranteed performance in every production environment. Indirect prompt injection can look different from the direct jailbreak examples used in testing. Performance may also change when the model receives tool results, retrieved documents, application-specific formatting, multiple languages, or benign instructions that resemble attack text.
Threshold selection is therefore important. A lower threshold may catch more attacks but also flag more legitimate content. A higher threshold may reduce disruption to users while allowing more malicious prompts through. Developers should calibrate the threshold against representative traffic and track false positives, false negatives, language differences, and changes in attack patterns.
Deployment and integration
Prompt Guard 2 86M can be loaded locally with Hugging Face Transformers, using either a text-classification pipeline or the lower-level AutoTokenizer and AutoModelForSequenceClassification classes. The supplied research identifies local and self-hosted Transformers inference as the supported deployment approach.
A typical placement is immediately before a trust-boundary transition. The detector can inspect a user's prompt before it reaches the main LLM, scan retrieved passages before they are inserted into a prompt, examine a tool result before an agent uses it, or screen an uploaded text document before downstream processing. The classifier's result should be combined with other controls such as permission checks, output validation, tool restrictions, and human review for high-risk operations.
The model supports fine-tuning. Application-specific training can be useful when the input format, domain vocabulary, tool-use patterns, or benign-content distribution differs substantially from the data represented in the base model. Fine-tuning should be evaluated carefully because improving detection for one attack pattern can affect false positives on legitimate content.
Pricing and access
No official per-token hosted API price was identified for this exact model. Prompt Guard 2 86M is distributed as an open-weight model through Hugging Face, but access to the download is gated and requires acceptance of the displayed license terms.
Because the supplied information does not identify a hosted inference endpoint or recurring provider price, there is no verified API price to report. Operating cost depends on the infrastructure used for self-hosting, the number and length of segments scanned, and the application's throughput requirements. The model's 86-million-parameter size is relevant to deployment planning, but the supplied research does not specify a hardware requirement or guaranteed latency for every environment.
Reasoning, coding, and tool support
Prompt Guard 2 86M is not a reasoning model. It does not perform multi-step problem solving for users, write software, or generate explanations of its classifications as a conversational assistant. Its classification decision can be used by a reasoning or coding system, but those capabilities belong to the surrounding application or another model.
The model has no native tool or function-calling capability. It does not invoke tools, browse the web, execute code, or take actions. It can inspect text produced by tools or intended for tools, which makes it useful as a security checkpoint in an agent workflow, but the orchestration logic remains the application's responsibility.
Main strengths and limitations
Strengths
- It is specialized for prompt-injection and jailbreak screening rather than being a general-purpose model with an indirect safety use.
- Its 86-million-parameter architecture is substantially narrower in scope than a full conversational model, making it suitable for a dedicated classification role.
- It supports multilingual evaluation across eight reported languages.
- It can be deployed locally or self-hosted with Transformers, which is useful when application data should not be sent to a hosted model.
- Logits are available, allowing developers to calibrate thresholds and build more nuanced review policies.
- Fine-tuning is supported for application-specific traffic and attack patterns.
Limitations
- The 512-token context window requires chunking for long documents, retrieved content, and tool outputs.
- The binary output does not distinguish prompt injections from jailbreaks, so applications that need separate reporting must add their own logic or use another classifier.
- It detects prompt-attack intent; it is not a complete content-safety system and does not independently judge every other type of harmful material.
- Reported benchmark results may not predict performance on indirect injections, unusual formats, new attack methods, or traffic in languages and domains unlike the evaluation data.
- It is not a conversational, coding, reasoning, image, audio, or video model.
- No official hosted API pricing or universal latency guarantee is provided in the supplied research.
When to choose Llama Prompt Guard 2 86M
Choose Prompt Guard 2 86M when the main requirement is screening text for prompt injections or jailbreak attempts before that text reaches a larger model or an agent action. It is particularly appropriate for multilingual applications, retrieval-augmented systems, document-processing pipelines, tool-output inspection, and deployments that favor local or self-hosted inference.
It is also a reasonable fit when a team wants a dedicated detector that can be thresholded, fine-tuned, and placed at several points in a security pipeline. Its narrow purpose can be an advantage over asking a general-purpose LLM to judge whether every input is malicious, especially when predictable classification behavior and local deployment are more important than open-ended explanations.
Another type of option may be more appropriate when the application needs a long-context safety analysis, detailed reasons for a decision, broad content moderation across multiple harm categories, or a general model that can both inspect content and carry out a complex task. A general-purpose LLM may be better for explanation and flexible reasoning, while a broader moderation system may be preferable when the policy covers more than prompt attacks. Prompt Guard 2 should be treated as one layer in a defense-in-depth design, not as a guarantee that an agent or LLM application is secure.

