Build1 publisher3 min readPublished
A well-known prompt-injection classifier caught 6 of 629 buried attacks at its default 0.5 cutoff
Scores in the buried-injections benchmark show a well-known prompt-injection classifier catching 6 of 629 hidden attacks at its 0.5 default and 621 at 0.003. A reviewer who re-ran the saved scores found its pooled cross-domain rates hide how detectors do on their weakest suite.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The classifier's attack scores ran about ten times its benign scores, so it ranked correctly, but the 0.5 cutoff sat two orders of magnitude above where those scores lived.
- The same benchmark post also carried the line "99% at a 2% false-alarm budget" as a headline result.
- On held-out suites the false-alarm rate came out above the 2% budget for eight of nine detectors; only the regex baseline stayed within it.
- Six of the nine detectors had a spread across folds at least as large as the pooled rate they reported.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision The shipped default has to be replaced. Because the scores already ranked attacks correctly, a cutoff set from the deployment's own benign traffic decides whether this detector works at all.
- constraint With under a hundred benign samples, a false-alarm budget moves in whole cases. A percentage budget then promises more precision than the data can deliver.
- exposure A team that cites one pooled catch rate in a security review can pass it with a detector that misses almost everything on one class of traffic.
- capability Because the repo publishes raw scores, anyone can re-score every detector under a different decision rule, such as worst-fold reporting, without running a model.
The 0.5 default is about 167 times the 0.003 threshold that worked on this data [1]. At 0.003 the classifier still missed 8 of the 629 attacks [2]. The figures come from the buried-injections benchmark. A reviewer writing on dev.to as pm25coder checked it against the repository rudratoshs/buried-injections at commit 232d0c13 [6].
In bench/at_budget.py, the function caught_at_budget sets its threshold in order [6][8]:
1. Compute `allowed = math.floor(budget * len(benign_scores))`. 2. Sort the benign scores from high to low and take the one at index `allowed` as the threshold. 3. Count an attack as caught only if it scores strictly above that threshold.
The saved results hold 97 benign scores per detector [7]. At a 2% budget, floor(0.02 x 97) is 1. The threshold is therefore the second-highest benign score, and one benign case may sit above it [3]. On any calibration split under 50 benign cases the allowance rounds to zero, and the threshold becomes the highest benign score [4].
A cross-domain wrapper runs that rule on three suites' benign cases and measures on the fourth, rotating through all four [9]. The 2% is enforced only on the split the threshold was picked from [8][9]. "The budget is a selection rule; the held-out false-alarm rate is a measurement," the reviewer wrote [11]. According to the review, the post's headline line puts a cross-domain catch rate beside an in-sample budget [5]. "Each number is honest; the sentence is the problem, because it is the sentence that gets screenshotted," the reviewer wrote [12].
The wrapper also adds the four held-out folds into one total [13]. The reviewer argued that a pooled rate shows the analysis ran on four suites without showing what the detector did where it was weakest [16]. A second detector caught 17%, 36%, 59% and 100% across the same suites and pooled to 51% [15]. That is an 83-point spread [6]. For any pooled figure here to describe a deployment, the deployment's traffic would have to look like the four-suite blend. Within one benchmark, changing which suite was held out moved the main detector's catch rate by 99 points [5].
The benchmark deserves credit for being checkable. The reviewer ported both functions, ran them over the repo's saved scores, and reproduced every number it saved for all nine detectors [17]. I would like more benchmarks to be this easy to audit. The post had asked for exactly this: "if you find a flaw in the methodology, I genuinely want to know," its author wrote. According to the reviewer, he has since said publicly he is fixing all of it [18].
For a detector sitting in front of tool output, I would rather read four fold rates than one pooled rate. The worst fold has its own problem. The reviewer wrote that leading with the minimum fold needs one more step, because that fold is, by construction, the smallest numerator [19].
What to watch
- Whether the author's promised fix changes the headline to per-fold or worst-fold catch rates, with held-out false-alarm rates in place of the in-sample budget.
- Per-suite benign counts in the repo: any calibration split under 50 cases means its 2% budget permits zero false alarms.
- Whether the classifier's maintainers change the 0.5 default or publish calibration guidance for tool-output traffic.