Build1 publisher2 min readPublished
Every open-weight model KaliBench tested fell below 42% on exact-command accuracy
KaliBench, 8,504 query-command pairs across 1,642 Kali tools, found no open-weight model among 24 configurations topped 42% exact-command accuracy. Scores rise when the model is handed the tool name, but production analysts describe only intent.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Existing agent benchmarks measure knowledge about tools or whether an agent finishes an attack end to end, and neither checks the step of turning intent into a correct invocation.
- The pairs cover 23 capability dimensions such as network scanning, password cracking and web exploitation, across five phases from reconnaissance to reporting.
- Scoring is at the argument level: a command fails for the wrong tool, a missing required flag, a semantics-changing flag reorder, or a bad value binding such as -p 80 where the tool needs -p80.
- KaliBench does not ask an agent to complete a penetration test; it asks for the correct command for a specific intent and verifies it without runtime execution.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability The runtime-free reward lets teams train security agents with PPO or DPO without ever running agent-written commands against a live network or target.
- exposure Operators are exposed because an agent can pick the right tool and return plausible-looking output that still will not execute, so a confident answer can be worthless in the field.
- constraint Even handed the correct tool name, models struggle to build the invocation, so tool selection and argument construction are two separate problems to fix, not one.
The claim under the benchmark is narrow: security agents grasp what you want but miss the exact syntax, flag bindings and argument order a CLI expects. [1]
Testing that is awkward, because you cannot run arbitrary Kali commands in a harness. Many need root, a network, or a live target; some are destructive; some trip intrusion alerts. [6] Running those to see how the agent did is not a sustainable test design. So KaliBench grades commands with a four-stage pipeline. It canonicalizes each command into a standard form, resolving aliases and normalizing flag order and whitespace. A separate model then checks whether the generated command is semantically equivalent to the ground truth. Commands judged safe run in a sandboxed terminal to confirm exit codes and output structure. Annotators settle the cases the automated steps leave ambiguous. [7]
The grader turns each command into a scalar reward: 1.0 for an exact match after canonicalization, 0.8 for a confirmed semantic equivalent, 0.5 for the right tool with wrong arguments, 0.0 for a wrong tool or a command that will not run. [13] Exact-command accuracy, the headline measure, counts only the 1.0 bucket. The 0.8 band is there because a tool usually has more than one valid invocation, and the strict score ignores every one of them by design.
The ceiling comes from the unrestricted setting, where the model picks the tool and builds the full command from an analyst's words. [10] Those were open-weight models, both general-purpose and security-focused. For that number to describe your own agents, your models have to sit in that class, and correct has to mean KaliBench's canonical command rather than any invocation that would also work. The published writeup breaks off as it begins to report how models fine-tuned on this reward performed. [15]
What to watch
- Whether the fine-tuned-model results cut from this writeup actually beat the 42% unrestricted ceiling.
- Whether closed frontier models, absent from the 24 open-weight configurations, clear the same bar.
- Whether security teams adopt the runtime-free reward to train agents with PPO or DPO instead of live execution.