Build1 publisher2 min readPublished
Artificial Analysis retries a provider safety error ten times before scoring the attempt zero
Artificial Analysis now reports how often a provider or model refuses coding-benchmark work on safety grounds. For a buyer the useful part is whether the agent stopped there or switched models and finished.
The Engineer · Build desk

What happened
- Artificial Analysis added safety-refusal reporting to its Coding Agent Index on September 18th, announcing it on X alongside version 1.5 of its coding-agent benchmark.
- It defines a safety refusal as a provider or model declining to start or continue a task on safety grounds, and splits those events into blocked attempts and fallbacks.
- The index runs 303 tasks across three evaluations, and the suite includes vulnerability discovery and exploitation work that can activate safeguards inside a controlled evaluation.
- The new chart reports refusals beside capability, cost, token use and execution time instead of folding every unsuccessful run into the same pass-rate number.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Anyone choosing between two agents with close scores has a new question to ask, because part of the score may have been earned by a fallback model weaker than the one they thought they were buying.
- exposure Security review and system administration are the workloads where legitimate work most resembles harmful activity, so they carry refusal risk that an aggregate coding score hides.
- constraint Refusals cannot explain a wide ranking gap, since the loss is capped at the blocked rate, so a buyer staring at a big score difference should be looking at capability.
Artificial Analysis scores a fallback attempt normally [7], and it says the fallback model may be less capable than the first choice [8]. So for any agent that routes around a refusal, the capability score is a blend. Some tasks were completed by the model you picked. The rest were completed by whatever the harness reached for next.
The retry rules are asymmetric. A provider safety error gets up to ten retries, and the attempt scores zero only if every retry is blocked [5]. A model refusal that ends an attempt scores zero right away, with no retry [6]. One zero can therefore sit behind as many as eleven calls to the provider [22]. The published rate counts retained, scored attempts and drops superseded retries [9], so it is a per-attempt figure. Your own error dashboard counts every call. The two numbers will not match.
A blocked rate of 2% could have cost an agent no more than two index points, Artificial Analysis says, because every blocked attempt scores zero [11]. The bound is proportional, so a 5% blocked rate caps the loss at five points [19]. The index runs 303 tasks across three evaluations [12]. If those three are equal in size, 2% is about six refused attempts [20].
For the published rate to predict what you will see, your task mix has to resemble the index's [13]. The overall rate gives each of the three component benchmarks equal weight [10]. A refusal in the smallest evaluation therefore moves the published figure more than a refusal in the largest. Under equal sizes, one refusal is worth about a third of a percentage point [21].
Micah Hill-Smith, now CEO, and George Cameron, now chief product officer, co-founded Artificial Analysis [17]. Hill-Smith has described the origin: while building a legal-research tool, he found different models worked better at different steps, with no dependable way to compare quality, speed and price [18]. Refusal reporting extends that comparison from model quality to provider policy.
The metric does not determine whether a provider's policy is appropriate, Artificial Analysis says; it measures the practical effect of that policy within a defined set of tasks [15]. The company also says buyers still have to decide whether a lower refusal rate reflects greater usefulness, looser safeguards or a harness that handles provider restrictions more effectively [16]. Two of those three readings describe engineering a buyer can test. The third describes how much safeguard the vendor relaxed, and the refusal rate on its own does not tell them apart.
What to watch
- Whether Artificial Analysis publishes blocked and fallback rates per agent for the security component rather than only the composite.
- Whether providers loosen filters on tasks that resemble the benchmark once refusal rates are public and comparable.
- Whether the index's task count or component sizes change in later versions, which would move the weight a single refusal carries.