Build1 publisher3 min readPublished
Renaming agent tools shifts attack-success scores by up to 13 points in Penn State tests
Penn State researchers moved agent attack-success rates by up to 13.21 points by renaming tools, with the attack, task and policy held fixed. The size of the shift varied by model and benchmark, so a robustness score only describes the tool names and schemas it was measured on.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- On MCPTox, giving a neutral tool an explicit threat-related name cut GPT-5-mini's attack success by 11.00 points but Claude Haiku 4.5's by only 4.11.
- A threat-neutral replacement name matched on token count, length and casing reproduced 8.54 of GPT-5-mini's 11.00-point MCPTox shift.
- On AgentDojo, threat-related wording moved GPT-4o-mini's attack success by 0.50 points while benign utility on tasks needing that tool fell 5.36 points.
- A dev.to write-up of the paper says enterprises compare models and choose defence mechanisms using these attack-success scores.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The same edit moved two models by amounts 6.89 points apart, so an ASR gap on one benchmark cannot be read as a robustness gap on a buyer's own tool schemas.
- cost A defence change judged only on attack success can pass review while costing legitimate task completion, so guardrail evaluations need the utility figure next to ASR.
- decision Anyone who wants a number that transfers has to re-run the attack suite against the production agent's own tool names, JSON keys and observation formats.
The control experiment tests the obvious explanation. A model that sees "delete" or "purge" in a tool name might simply get cautious. If threat words were the whole cause, a neutral name of the same shape would move nothing. According to a dev.to write-up of the Penn State paper, it moved most of the score [1][5]. Only 2.46 of GPT-5-mini's 11.00 MCPTox points are left for the threat wording itself [9]. About 77.6 percent of the shift came from changing the string at all [10].
The write-up concludes: "The representation itself, not semantic content, drives the variance." [7] It then attributes the effect to tokenization boundaries interacting with positional encodings, attention masks and layer activations [14]. The control supports the first sentence for one model on one benchmark. The second is the writer's hypothesis. The reported experiments change inputs and measure outputs [5].
For model selection, a format effect matters when it differs between models. A constant offset would leave rankings intact. The reported offsets differ. On Agent Security Bench, neutral tool names raised committed attack success by 11.67 points on GPT-5-mini and 13.21 on Claude Haiku 4.5 [2]. On MCPTox the two models ended up 6.89 points apart on the same edit, with GPT-5-mini the more sensitive [3][11]. Which model is more format-sensitive depends on the benchmark. The write-up does not report baseline attack-success rates, so this material cannot show whether any ranking actually flipped.
ASR is the share of attempts in which the agent completes the harmful action [6]. Those attempts run against one benchmark's tools, in that benchmark's names, JSON keys and prose. For a published score to transfer, a production agent's tools would have to look to the model the way the benchmark's do. Teams that spend a week arguing over tool naming conventions now have a security reason to keep arguing. The write-up says enterprises compare models on these scores [15]. It argues that "A 5-point ASR difference might just be representation noise." [8] The ASB renames moved the score by more than twice that [13].
The AgentDojo case shows how this goes wrong in defence evaluation. The wording change barely registered in attack success. It showed up in benign utility instead [4]. A review that scored the change on ASR alone would have approved it. The write-up extends the point to guardrails, warning that a prompt filter tuned on JSON schemas "might fail on prose observations." [16]
The design is sound. Six things stay fixed, including the harmful action, the security policy and the evaluation criteria. Only the agent-visible representation changes [1]. The 11-to-13-point figure is the top of the reported spread. Across the cases in the write-up, shifts ran from 0.50 points to 13.21 [12]. The models were GPT-5-mini, Claude Haiku 4.5 and GPT-4o-mini [2][4].
What to watch
- Publication of baseline ASRs per representation, which would show whether model rankings on ASB or MCPTox flip under tool renames.
- Results on larger models than GPT-5-mini, Claude Haiku 4.5 and GPT-4o-mini, to see whether the format effect shrinks or grows.
- Whether ASB, MCPTox or AgentDojo add multiple naming and schema variants to how they score attack success.