Skip to content

Build1 publisher3 min readPublished

Renaming agent tools shifts attack-success scores by up to 13 points in Penn State tests

Penn State researchers moved agent attack-success rates by up to 13.21 points by renaming tools, with the attack, task and policy held fixed. The size of the shift varied by model and benchmark, so a robustness score only describes the tool names and schemas it was measured on.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Renaming agent tools shifts attack-success scores by up to 13 points in Penn State tests
Generated illustration

What happened

  • On MCPTox, giving a neutral tool an explicit threat-related name cut GPT-5-mini's attack success by 11.00 points but Claude Haiku 4.5's by only 4.11.
  • A threat-neutral replacement name matched on token count, length and casing reproduced 8.54 of GPT-5-mini's 11.00-point MCPTox shift.
  • On AgentDojo, threat-related wording moved GPT-4o-mini's attack success by 0.50 points while benign utility on tasks needing that tool fell 5.36 points.
  • A dev.to write-up of the paper says enterprises compare models and choose defence mechanisms using these attack-success scores.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The same edit moved two models by amounts 6.89 points apart, so an ASR gap on one benchmark cannot be read as a robustness gap on a buyer's own tool schemas.
  • cost A defence change judged only on attack success can pass review while costing legitimate task completion, so guardrail evaluations need the utility figure next to ASR.
  • decision Anyone who wants a number that transfers has to re-run the attack suite against the production agent's own tool names, JSON keys and observation formats.

The control experiment tests the obvious explanation. A model that sees "delete" or "purge" in a tool name might simply get cautious. If threat words were the whole cause, a neutral name of the same shape would move nothing. According to a dev.to write-up of the Penn State paper, it moved most of the score [1][5]. Only 2.46 of GPT-5-mini's 11.00 MCPTox points are left for the threat wording itself [9]. About 77.6 percent of the shift came from changing the string at all [10].

The write-up concludes: "The representation itself, not semantic content, drives the variance." [7] It then attributes the effect to tokenization boundaries interacting with positional encodings, attention masks and layer activations [14]. The control supports the first sentence for one model on one benchmark. The second is the writer's hypothesis. The reported experiments change inputs and measure outputs [5].

For model selection, a format effect matters when it differs between models. A constant offset would leave rankings intact. The reported offsets differ. On Agent Security Bench, neutral tool names raised committed attack success by 11.67 points on GPT-5-mini and 13.21 on Claude Haiku 4.5 [2]. On MCPTox the two models ended up 6.89 points apart on the same edit, with GPT-5-mini the more sensitive [3][11]. Which model is more format-sensitive depends on the benchmark. The write-up does not report baseline attack-success rates, so this material cannot show whether any ranking actually flipped.

ASR is the share of attempts in which the agent completes the harmful action [6]. Those attempts run against one benchmark's tools, in that benchmark's names, JSON keys and prose. For a published score to transfer, a production agent's tools would have to look to the model the way the benchmark's do. Teams that spend a week arguing over tool naming conventions now have a security reason to keep arguing. The write-up says enterprises compare models on these scores [15]. It argues that "A 5-point ASR difference might just be representation noise." [8] The ASB renames moved the score by more than twice that [13].

The AgentDojo case shows how this goes wrong in defence evaluation. The wording change barely registered in attack success. It showed up in benign utility instead [4]. A review that scored the change on ASR alone would have approved it. The write-up extends the point to guardrails, warning that a prompt filter tuned on JSON schemas "might fail on prose observations." [16]

The design is sound. Six things stay fixed, including the harmful action, the security policy and the evaluation criteria. Only the agent-visible representation changes [1]. The 11-to-13-point figure is the top of the reported spread. Across the cases in the write-up, shifts ran from 0.50 points to 13.21 [12]. The models were GPT-5-mini, Claude Haiku 4.5 and GPT-4o-mini [2][4].

What to watch

  • Publication of baseline ASRs per representation, which would show whether model rankings on ASB or MCPTox flip under tool renames.
  • Results on larger models than GPT-5-mini, Claude Haiku 4.5 and GPT-4o-mini, to see whether the format effect shrinks or grows.
  • Whether ASB, MCPTox or AgentDojo add multiple naming and schema variants to how they score attack success.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories