Security1 distinct publisher2 min readPublished
Cisco Talos reports that turning reasoning effort up often cost more without scoring better, and sometimes scored worse. The number that should drive procurement is the worst run, not the median.
The Watch · Security desk
Compiled by The WatchSomething wrong?How this is made
The scoring rules do more work here than the findings. A refusal or a malformed report never lowers a condition's score: the panel is discounted after a limited number of retries, so the run leaves the quality measurement entirely and the wasted attempt turns up in cost and wall time instead [7][11][13]. Quality was also smoothed twice before anyone looked at it, first as a mean across four personas in a panel, then as a median across that condition's panels [8][6]. Two layers of averaging is a reasonable way to suppress noise, and it is also a reasonable way to hide the single bad run that Talos itself flags as the thing a SOC cannot absorb [10].
Both economic metrics divide everything spent, failures and retries included, by the number of complete usable panels [11][13]. An unreliable condition therefore pays twice for the same fault: the discarded attempt stays in the numerator while the denominator shrinks [2]. That is the closest thing in the design to a reliability price, and it is the one a rota of analysts actually feels, because refusals and format breakage consume a shift whether or not they produce a score.
The arithmetic of adopting this is not small. Five rounds across 66 conditions is 330 planned panels, and at four personas each that is 1,320 planned agent investigations before a single retry [1]. A team that wants its own ranking rather than someone else's is budgeting four figures of agent runs per evaluation pass.
What gets compared is also model plus harness, not model alone: Anthropic conditions ran in Claude Code, OpenAI conditions in Codex [4]. Within a vendor, the reasoning-effort deltas are clean. Across vendors, the harness rides along in every number, which is defensible, since that is how these tools are bought and run.
One definition needs a decision before replication. Talos states downside consistency as the median score for the panel minus the lowest score in that panel [14], while scores aggregate from panel means up to a condition median [8], so a reproducing team has to choose which level the gap is measured at. And the score itself is agreement with one known answer on one synthetic corpus that reviewers were told might be real [3][5], picked because triage work rarely collapses into a single number [17]. Talos's stated output is the method, not a leaderboard [2], and given that effort reliably raises the bill without reliably raising the score [9], effort belongs in the same category as any other tunable that has to be measured per workflow.
Ranked by verification strength, evidence, and original report placement.
Using only common Unix command-line tools, reviewers had to decide whether a given dataset was real or synthetically generated. Each reviewer received an identical dataset. The dataset was synthetic, but reviewers were told it might be real.
Talos found reasoning effort was not a universal quality dial: more effort often cost more without improving the result, and in some cases more effort produced lower scores.
Talos chose the task because it required many of the same tools and analytic techniques used in typical incident triage and investigation, but unlike those scenarios could easily produce a single numeric score for comparison.
Cisco Talos tested 66 model and reasoning combinations from Anthropic and OpenAI on a tool-assisted log-review task.
Talos reports that instead of identifying a clear winner, it found a repeatable methodology that organizations can use in their own evaluations.
Reviewers investigated the logs using their native agent harnesses: Anthropic models used Claude Code, OpenAI models used Codex.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed first-party method, one task, no published per-condition data
The methodology is unusually well specified for a vendor blog: 66 named conditions, five planned rounds each, a four-persona panel with an all-valid-reports validity rule, explicit mean/median aggregation, a downside-consistency definition, wall-time and cost accounting that include failures and retries, and a corpus generator pinned to a frozen version with ground truth withheld. It also self-discloses its main limitation, that cost is a frozen public list-price proxy rather than incurred spend. What holds the score down: a single synthetic-data-detection task over one six-hour scenario, only five planned rounds per condition, no reported dispersion statistics or significance testing, only two hosted-model vendors in scope, and no per-condition results table in the supplied excerpt, which is truncated mid-discussion. Everything rests on one first-party source with no independent replication.
First-party run only; no third-party uptake shown
The one concrete adoption fact is Talos running and publishing the evaluation itself, at meaningful scale (66 conditions, an implied 330 planned panels and 1,320 persona investigations), using its own open-source EvidenceForge generator frozen at v1.12.0. That is genuine deployed usage of the harness inside one security research organization, and EvidenceForge being open source makes reuse possible. But the supplied material shows no other organization adopting the methodology, no SOC deployment outcome, no downstream tooling, and no procurement decision attributed to it, so adoption stays near the floor rather than being unmeasurable.
Cautious framing, but guidance generalizes past one task
Talos deliberately deflates the usual benchmark story: it refuses to name a winner, warns that chasing the top median score could have severe negative consequences, and discloses that its cost figures are frozen list prices. That restraint pulls the gap toward zero or below. Pushing it slightly positive is the breadth of the portable conclusions - that reasoning effort is not a universal quality dial and that consistency should be a major procurement factor - drawn from one synthetic-log task, one scenario, five rounds per condition and two hosted vendors, with no per-condition data shown in the supplied excerpt and no independent replication. The overstatement is one of scope, not of substance.
Security vendor benchmarking third-party model vendors, promoting own tooling
Cisco Talos is the research arm of a vendor that sells security operations products, and the post both evaluates third-party model providers (Anthropic, OpenAI) and showcases Talos's own open-source EvidenceForge generator and evaluation methodology. The framing that model choice is a complex multi-variable balancing act is congenial to vendors selling orchestration and SOC tooling. Mitigating factors are substantial: no Cisco model or product is scored, no winner is declared, cost is normalized to public list prices explicitly for cross-provider comparability, and limitations are disclosed. There is no conflict-of-interest statement in the supplied text, so the incentive is real but moderate rather than acute.
Method well documented, corroboration absent
Confidence is moderate. What Talos says it did and concluded is unambiguous and internally consistent, so claims about the study's design and stated findings are safe. Confidence in the findings themselves is limited by having exactly one publisher in the cluster, a truncated body that omits the per-condition results, a single task and scenario, five planned rounds per condition, list-price cost proxies, and zero third-party replication or adoption evidence.
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
leadership
Plan for both: AI reshapes the economy and most of this wave's capital is lost1 distinct publisher
security
Uber ships ADR, and hands agent-security vendors a number to be measured against1 distinct publisher
build
OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit3 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.