Build1 publisher3 min readPublished
Red Hat's 200M-parameter classifier nearly ties a 35B LLM judge on prompt injection
Red Hat's 200M-parameter classifier scored 89.01% on prompt injection to a 35B LLM judge's 89.31%, answering in a median 54.1 ms against 312.5 ms. For injection screening, a large judge in the request path now has to justify nearly six times the latency for 0.3 points.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Red Hat's AI Safety team ran nine guardrail setups, covering classifiers, LLM judges and decision models, through NVIDIA's NeMo Guardrails on prompt-injection and content-safety tests.
- A policy tuned for the Laya model lifted it from 57.87% to 75.20% on content safety, and the same policy cut Jev from 86.20% to 82.53%.
- Red Hat plans to ship both of its classifiers as the default guardrail configurations in OpenShift AI 3.6.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams that put an LLM judge in front of every request for injection screening are trading about 258 ms of median latency for 0.3 points of accuracy, on Red Hat's numbers.
- constraint No small classifier in the test came within five points of the leader on a broad content policy, so a guardrail stack using small models has to split by risk category and keep a larger model where the policy is wide.
- cost Policy rewrites moved single models by 15 to 17 points, far more than the gap at the top, so each adopter has to re-run candidate judges against its own policy before a published rank means anything.
- exposure OpenShift AI 3.6 users who keep the defaults will get a content-safety classifier that trailed the leader by about six points in Red Hat's own benchmark.
Dividing 35 billion parameters by 200 million gives 175x, and that ratio overstates the compute gap [26][9]. Qwen3.6-35B uses a mixture-of-experts design, so only about 3 billion of its parameters are used on any given token [9]. Against a 200M classifier, the ratio of active weights is about 15x [26]. Latency is what an operator pays on each request. Red Hat's main table puts DeBERTa at a 54.1 ms median and Qwen at 312.5 ms. The gap is about 258 ms, or 5.8x [1][20][24].
For a check that runs on every request, I'd use the classifier for prompt injection. Paying 258 ms per call for 0.3 points of accuracy makes sense only if your traffic is much harder than Red Hat's test set [25][24]. For the number to transfer, your injection attempts have to look like the ones Red Hat tested, and your serving stack has to produce a similar latency spread [4].
Content safety is a wider task. Red Hat's policy took in violence, prejudice, profanity, sexual content and illegal activity, along with role play and similar pretexts [15]. Jev and DiffusionGemma beat the 125M Granite Guardian classifier on it by more than five points, and the gap to Jev was 5.93 [15][27]. Granite was still the fastest entry, at a 33.2 ms median [11]. The authors acknowledged that the result points to a need for better small predictive models in that category [13]. Credit to them: Red Hat plans to ship Granite as a default in OpenShift AI 3.6, and it still published a table in which Granite finishes sixth [12][11].
Decision models were the third approach in the test. Jev takes the application's state and a set of typed questions. It returns typed answers, such as a probability between 0 and 1 [14]. The pitch is LLM flexibility without paying to generate tokens the application never uses [14]. In Red Hat's setup the saving did not show up as speed. Qwen posted lower median latency than Jev on both benchmarks [16].
Jev's accuracy lead is also thin. DiffusionGemma, an open Jev-style model served through vLLM, came within 0.67 points of it on content safety [10][18]. It also beat Jev on prompt injection, 87.72% to 86.35% [18]. NVIDIA's 4B Nemotron-3.5-Content-Safety trailed Jev by 1.13 points on content safety while responding faster [17]. TypeSafe's launch numbers for Jev came from evaluations it designed and ran itself [6].
Policy wording moved individual scores by up to 17 points [22]. When Red Hat swapped out NVIDIA's default risk definitions for its own, Nemotron's prompt-injection accuracy went from 69.37% to 84.84% [2]. That 15.47-point move is more than fifty times the gap between DeBERTa and Qwen [21][23]. Laya is an open decision model of about 421M parameters that Red Hat ran on a laptop CPU [19]. It scored 57.87% on content safety under the original policy and 75.20% under one tuned for it [3]. The same tuned policy cut Jev from 86.20% to 82.53% [3]. Each row in the table is a model paired with a policy, and the policy that gained Laya 17.33 points cost Jev 3.67 [22].
What to watch
- A small content-safety classifier from Red Hat or anyone else that closes the roughly six-point gap to Jev and DiffusionGemma on a broad policy.
- Independent reruns of the nine configurations with other policies and test sets, given that policy wording alone moved some models 15 to 17 points.
- Whether OpenShift AI 3.6 lets operators swap the shipped classifier defaults for a decision model or LLM judge per risk category.