Build1 distinct publisher3 min readUpdated
A football similarity space scored 0.695 for possession contamination, and the obvious one-line normaliser only got it to 0.572. Subtracting a fitted line instead got 0.450.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The arithmetic that breaks the obvious fix fits on one line. A per-90 count behaves like a + b * possession, and the intercept is real, because a player receives the ball and plays it whether or not his team dominates [11]. Divide by possession and you get a/possession + b, where the first term is a hyperbola that falls steeply as possession rises, so the operation does not remove the possession signal, it substitutes one of the opposite sign [12]. Division is correct only when the relationship passes through the origin, and according to the author it almost never does [13].
Only a scored defect exposes that. Uncorrected, the space read 0.695 against a chance baseline of 0.000 [7][10]; dividing brought it to 0.572, and the residual method brought it to 0.450 [8][9]. That is 17.7 per cent of the contamination removed by division and 35.3 per cent by residualising [1], which puts the ten-minute version at almost exactly half of what was available to remove [2]. On its own, 0.572 would have passed as a fix.
Recomputing pays elsewhere too. The multiplier quoted for how much more often a Barcelona player touches the ball is 1.74, while dividing the two possession shares as printed gives 1.76 [5][6]. Rounding, not error, but you only see it by redoing the division yourself.
What survives the better operation deserves more attention than it gets. At 0.450 the space still carries 65 per cent of its original contamination [3], and the writeup does not account for the remainder. Role coherence moved from 76.3 per cent to 76.8 per cent [15], so the correction cost nothing in the property the space was built for. Free and harmless is not the same as sufficient.
The earlier attempt is the same failure with a different denominator. The theory was that a ratio cannot scale with exposure the way a count does, so ratios should be immune by construction [19]. Passes per 90 has 24 per cent of its variance explained by which club a player is at; pass completion, the ratio brought in to replace it, has 21 per cent [16]. The substitution removed about an eighth of the problem [4]. Completion ratios inherit team style because a dominant side offers easier passes, while shape metrics stay clean, with final-third touches at 3.2 per cent and shots per 90 at 3.6 per cent [16][17].
The part that carries over to requests per user or errors per deploy [20] is the choice of diagnostic. Contamination was first tracked as the share of a player's eight nearest neighbours who were his own club teammates: 3.0 per cent against a 1.3 per cent chance baseline [18], about 2.3 times random [7], which reads as a blemish not worth an evening. The same vectors, scored on possession instead, read 0.695 against 0.000. One defect, two verdicts, and the normaliser written next depends on which measure happened to exist first.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The author built a vector space of 1,419 footballers from 5.3 million events across four leagues, with each player represented as 17 dimensions (passes per 90, progressive carries per 90, share of aerial duels won and so on), each converted to a percentile against every other player, and cosine similarity between vectors read as playing alike.
Position is never given to the model, and the fact that the eight players nearest a centre back are eight centre backs is the evidence that the space captures playing style.
Gerard Pique's nearest neighbour in the space was Jeremy Mathieu, his own centre-back partner at Barcelona, and Sergio Busquets sat next to two other holding midfielders from Spanish possession sides.
Barcelona had 67 per cent of the ball that season and Carpi had 38 per cent, so a Barcelona player touches it 1.74 times as often as a Carpi player and every per-90 count partly measures his employer.
To measure the defect itself the author used the correlation between a player's team possession and the mean team possession of his eight nearest neighbours.
With no correction, the possession contamination score was 0.695.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific self-reported numbers, no external verification
The methodological core is checkable on its face: the affine-versus-proportional algebra is mathematically explicit, the residualisation is shown in two lines of numpy, and the outcome is reported with a defined metric, a chance baseline and three comparable scores plus a control (role coherence). Against that, everything numeric is a single self-report from one author on one dataset, with no code, data, variance-decomposition method, uncertainty estimates or independent replication in the cluster, and one small internal arithmetic slip (1.74x versus the 1.76x implied by the printed possession shares).
No third-party adoption evidence
The only adoption-shaped fact in the cluster is the author deploying his own hobby app behind a domain. There are no users, downloads, third-party deployments, benchmark entries, or reports of anyone else applying the residualisation approach, so adoption cannot be scored without inventing facts.
Broadly aligned, slightly understated
Framing tracks the evidence closely and errs conservative: the title and dek foreground the failed correction rather than the win, the adopted result is explicitly called a partial fix (0.450 against a 0.000 baseline), the residual tactical-style confound is written down as a limitation, and the Umtiti coincidence is labelled a coincidence rather than validation. The generalisation to requests per user and errors per deploy reaches beyond what is measured, which pulls the gap back toward zero.
Mild self-promotional incentive, no disclosed commercial stake
This is a self-published developer-platform post: the author benefits reputationally from a clever methodological result, and the piece closes by promoting a tooling workflow (written with Claude Code, driven remotely from a phone), which is a soft product endorsement. No vendor funding, employer stake, sponsorship or commercial product of the author's is disclosed or evident, and the post's willingness to publish its own failed correction and unresolved limitation runs against pure boosterism, so incentive pressure reads low-to-moderate.
Moderate: sound reasoning, unverified measurements
Confidence in the reasoning is high - the affine-intercept argument and the proxy-blindness point stand on their own logic and are the parts most likely to transfer. Confidence in the specific magnitudes is much lower: one publisher, one author, one undisclosed dataset, no replication, no error bars, and explanatory mechanisms (why completion ratios inherit team style) asserted rather than tested. The blended score reflects a credible but unaudited single-source account.
build
The harness, not the model: 250 scores were the acceptance spec for a solo MusicXML editor1 distinct publisher
build
Make the spec fail the build: a solo dev's log of docs-versus-code drift1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
build
Superpowers makes spec-driven work a precondition, then ships it to twelve harnesses1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 23, 2026