Build1 publisher3 min readPublished
One threshold change took Meta's Prompt Guard 2 from 1% to 99% of buried injection attacks
Meta's Prompt Guard 2 caught 6 of 629 buried attacks at a 0.5 cutoff and 621 at 0.003, in a benchmark posted on dev.to. Health checks pass the same at both settings, so only attacks sent through the deployed threshold show which one a team runs.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The author buried 629 injection attacks from the AgentDojo benchmark inside ordinary tool output, such as bills, emails and web pages, and ran 10 open-source detectors over them.
- At the conventional 0.5 cutoff, attack and benign scores alike fell far below the line, so Prompt Guard 2 returned the same pass verdict on every request.
- The 0.003 cutoff was calibrated to wrongly flag at most 2% of normal traffic, measured on an AgentDojo domain the model was never tuned on.
- The author called most of the other detectors' results "the boring kind of bad": weak detection, or alarms on harmless input.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure An agent running this detector at a conventional cutoff stays open to injected tool output for as long as nobody sends a real attack through the production path to check.
- decision Setting the cutoff becomes a calibration job on each team's own attacks and benign traffic, since the working value here sat about 167 times below the conventional one.
- contradiction The post calls 0.5 both a convention users apply and a default the model shipped with; the answer decides whether the fix belongs in Meta's documentation or in every integration.
A detector like Prompt Guard 2 has two parts. One is a model that emits a probability that the input is malicious. The other is a cutoff in the calling code that turns that probability into a block [3]. In the post's example, a buried attack scored 0.009 and benign text scored 0.0008 [4]. The ratio is about 11 to 1 [1], so the model ranked the attack correctly. "The detector was producing different scores. The control just wasn't doing anything with the difference," the author wrote [8].
A health probe only sees the model. In this configuration the detector loaded, returned scores, passed health checks and logged clean [15]. Liveness, latency and a schema check on the score field all exercise the model. A random number generator would pass the schema check too. "Every health check passes at 1% exactly as it does at 99%," the author wrote [9].
The post contradicts itself on where the 0.5 came from. The author wrote that "0.5 isn't some evil value Meta hard-coded" and described it as the threshold anyone gets by treating a probability classifier the normal way [10]. Later the same post calls it "a config value the model shipped with" and says "the shipped default was wrong by ~100x" [11]. One passage puts the error at about 50x, another at roughly two orders of magnitude [12]. Each figure fits a different quantity. The 0.5 cutoff is about 56 times the example attack score [2]. The move from 0.5 to 0.003 is a factor of about 167 [3]. The post does not quote Meta's documentation on an intended threshold.
The 99% comes from one benchmark [1]. The author cautions readers against taking the 99% to mean that Prompt Guard 2 solves prompt injection [13]. For the number to carry over, a deployment's injected text has to score in the same band. Its benign tool output also has to stay under 0.003 at the rate measured here. In the example, the new cutoff sits 3 times below the attack score and 3.75 times above the benign one. In absolute terms those gaps are 0.006 and 0.0022 [5]. Eight attacks still got through [4].
In my view, the guardrail check that counts runs end to end. It sends a fixed set of known attacks, buried the way the agent receives them, through the production path at the production threshold. It alerts when the block rate drops. It also needs a benign set, because a cutoff low enough to catch nearly everything can start blocking bills and emails. The benchmark got this part right. It included 97 benign tool outputs to catch detectors that alarm at everything [2].
What to watch
- Meta documentation or a model card stating a recommended Prompt Guard 2 threshold for text embedded in tool output, which would settle whether 0.5 is a shipped default.
- A replication on attacks from outside AgentDojo, testing whether a 0.003 cutoff holds near 99% detection at a 2% false-positive rate on other agents' tool output.
- Per-detector scores for the other nine detectors in the benchmark, to show whether any of them also fail at the threshold rather than in the model.