Build1 publisher3 min readPublished
Kendall's tau between backup time and archive size is -0.50, which leaves six of the nine plugins Pareto-optimal and makes the single number at the top of the board a weighting its author would be choosing for you.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Nine plugins make 36 unordered pairs. A Kendall's tau of -0.50 across those pairs works out to 27 discordant and 9 concordant, assuming no ties, so three quarters of the head-to-head comparisons flip when the ranking axis changes from backup duration to archive size [18]. The two orders are not loosely related. They are mostly reversed [4].
The cause is arithmetic in the CPU rather than a quirk of one vendor: compression spends time to save bytes, so the plugin that skips it finishes sooner and ships a fatter archive [5]. A wall-clock ranking is therefore partly a ranking of who compressed least. Six of the nine plugins are beaten on neither axis, and only three are beaten on both [6]. Arguing that the top line of your own leaderboard should not exist is an unusual move from second place, which is where the benchmark's author sits, by his own disclosure, as the seller of one of the nine plugins measured [1].
For the -0.50 to transfer to a benchmark someone hands you, the two axes on that board would have to trade against each other by construction, the way CPU time trades for bytes here. Where the axes move together, one number is a fair summary. The portable part is the question rather than the coefficient: which second metric was measured, and what weighting collapsed it into the headline figure.
SPEC has had a written position on this for years. Its Fair Use rules define a derived value as any number that is a function of a benchmark metric plus something else, permit composites outright, and prohibit presenting one as the benchmark's own metric [7]. The worked example is a gaming society building a composite from weighted subsets of two SPEC benchmarks: useful and interesting, says SPEC, but the weighting and the subsetting were the society's work, so calling it SPECgame would be a violation [8]. Two neighbouring rules do the enforcement. The basis for comparison has to be stated, meaning fastest among what, measured how, retrieved when [9]. Several benchmarks also require named metrics to be quoted alongside any headline number, so the flattering figure cannot travel alone [10].
The error bar had the same disclosure problem in a different coat. The board started with CI = 1.96 * (400 / sqrt(n)) over battles played, which gives +/- 351 at five battles and +/- 55 at 200 [11], and which assumes every battle is an independent observation [12]. One record had 24 of its 40 battles fail on the same fault reproduced from a single bug [13]. That is one fact seen 24 times. The fix treats outcomes sharing a root cause as perfectly correlated, so a cluster of any size counts once, with failures grouped by harness version, hosting tier and phase [14][16]. The code subtracts size minus one per cluster and floors the result at the number of distinct clusters seen [15]. Forty minus 23 is 17 effective observations [20], and the interval moves from +/- 124 to +/- 190 [17], about 53 percent wider [19].
Wider is the direction the evidence pushes, since the assumption being replaced was the flattering one. Both corrections cost the author his headline: one number becomes two axes plus a stated weighting, and a tight interval becomes a loose one that survives contact with a single bug.
Ranked by verification strength, evidence, and original report placement.
The benchmark's author discloses that he sells one of the nine plugins in the benchmark, and that it is currently second on the board.
The benchmark was built to answer which WordPress backup plugin is fastest: two plugins get byte-identical fixed snapshots on identical sites, backup and restore run in containers with no outbound network, each pairing does a warm-up and three measured runs, and results feed an Elo rating.
Every run produces two comparable outputs, how long the backup took and how big the archive came out, and ranking the field by each gives two orders.
Kendall's tau between the duration ranking and the archive-size ranking is -0.50, which the author describes as actively opposed rather than unrelated.
Across the whole field the plugins that finish fastest write the biggest archives, because compression costs CPU time and skipping it buys seconds and costs bytes.
Six of the nine plugins are Pareto-optimal, with no other plugin beating them on both speed and size at once; only three are beaten on both axes and can honestly be called worse.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One interested measurer, checkable arithmetic
Every measurement traces to a single harness run by a man who sells one of the entries, and nobody in our coverage has rerun it. What survives that is the arithmetic, which is unusually exposed: the formula, the eleven-line function and both intervals are printed, so 40 battles less a 23-observation cluster leaving 17, and 1.96 * (400 / sqrt(17)) landing on ±190, are checkable without trusting him. The SPEC material carrying the argument's authority is the weakest link — the derived-value section and the GamePerfMark ruling are recounted, not cited.
No uptake reported
Nothing in this reporting says who uses the benchmark or the plugins in it: no install counts, no downloads, no third party quoting the board. The Elo table is the author's own output, which tells us his harness ran and nothing about anyone acting on it.
Argues against its own best number
The strongest figure available to the author — 1442 against 1238 on the overall column, the one he says he would use if he were selling — is the one he refuses to call settled, because the intervals still overlap by 22 points. The correction he shipped widens his own uncertainty by roughly half and moves his rating not at all. Against that, borrowing SPEC's derived-value language lends a nine-plugin hobby harness more institutional weight than it has earned.
Seller grading the field he sells into
He sells one of the nine, it is second on the speed column and first overall, and the conflict is the first line of the post rather than a footnote at the bottom. Disclosure does not dissolve it: he picked the field, the snapshots, the hosting tier, the failure grouping and which column a reader meets first — and that last choice, the one he names as the whole argument, is still made on the reader's behalf.
Specific enough to dispute, single-sourced
The method claims are precise in the way that invites contradiction — a named formula, a printed function, a test asserting direction rather than magnitude — and the author's incentive is stated instead of inferred, so a reader knows where to discount. Holding it down: one account carries all of it, and the controls that matter most to a reader are asserted rather than shown. Byte-identical snapshots, network-isolated containers and six unnamed plugins are all things we would have to take his word for.
build
All-in-One WP Migration runs the attacker's stored SQL while rewriting URLs on restore1 publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 publisher
build
fal reports 35x MiniMax's own H3 endpoint after tuning the weights to its runtime1 publisher
security
Six bugs, one order of operations: Avada's zero-click chain is a same-day patch1 publisher
Publishers with included, body-backed reporting in this cluster.
1 article · September 7, 2026