Build1 distinct publisher3 min readUpdated
A per-asset Venn-Abers audit of seven crypto models found ETH claiming 93.8% and delivering 45.5%, while XRP delivered 93.1% at 64.9% stated. Opposite errors need opposite corrections.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Linearity is the whole mechanism, and it leaves no slack. Each term in the score is a probability multiplied by a fixed payoff [2], so a probability that is wrong by a factor enters the score wrong by the same factor [3]. Nothing downstream absorbs it, because the number that decides whether to fire is the number that decides how large.
Run the two failure modes through that arithmetic. ETH's top bin states 93.8% and delivers 45.5% [9], a ratio of about 2.06 [1], so a score built from that bin is close to double what the observed hit rate supports. XRP's [0.60, 0.70] bin runs the other way, 93.1% delivered against 64.9% stated [10], a factor of about 1.44 [2], which discards roughly 30% of the score before the gate evaluates it [3]. LTC sits at 1.40 in the same bin [11][7].
The counts decide which of those is the larger bill. ETH's bad bin holds eleven samples [9], 2.2% of the 500-row evaluation set [4]. The two understated bins hold 116 and 115 [10][11], 231 together [5], about twenty-one times as many observations [6]. Over-confidence mis-sizes a thin tail. Under-confidence declines a fat middle of trades that were good [22], and it leaves no losing fill behind to investigate.
Having both directions inside one universe is what rules out the cheap fix. Platt and temperature scaling apply a single monotone correction, so they cannot pull ETH's tail down and lift XRP's middle at the same time, and fitted globally they land on the average of two opposed errors [13]. The per-asset Venn-Abers wrapper works here because it takes each asset's own reliability curve without assuming a shape [14].
The provenance of the under-confidence is the part worth sitting with. min_data_in_leaf=400, num_leaves=8 and max_depth=3 were introduced specifically to stop terminal leaves reaching 1.0 and claiming certainty the model had not earned [15]. They did that, and they also produced a structural under-confidence pattern across the other six assets [16], while ETH still saturates in the rare cases where it is genuinely certain [17]. The leaf constraints did not remove a calibration problem. They swapped a loud one in one asset for a quiet one in six, and an aggregate accuracy figure would have reported a modest improvement and shown nothing else, since averaging is the operation that hides a bin [18].
LINK-USD is the loose end. It misses the acceptance gate on both axes, -41.5% on ECE and -1.2% on log-loss, against thresholds of 50% and 5% [19][20], and it is not short of data: 8,736 training rows, the same as every other asset [21], which puts 6,988 rows under the wrapper fit, the same 80% share the others got [6][8]. The author reads this as the LightGBM already being better calibrated for LINK, leaving less for the wrapper to take [23]. That is plausible and this test does not establish it. A wrapper that fails to improve a model has told you only that the wrapper did not help; distinguishing an already-flat reliability curve from a wrapper fitting noise takes LINK's own diagram.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
For each of seven live assets, an Inductive Venn-Abers wrapper was fit on the time-ordered older 80% of that model's training data, 6,988 rows, and evaluated against a 500-row uniform-random sample of the newer 20%, seed 42; the LightGBM models were reloaded from disk and left alone, with only the wrapper fit.
Expected Calibration Error and log-loss were measured before and after the wrapper.
ETH's worst reliability bin is [0.90, 1.00], with eleven samples, mean stated confidence 93.8% and empirical accuracy 45.5%; that is the bin where the EV gate fires hardest and position sizes are largest.
XRP in the [0.60, 0.70] bin shows 93.1% accuracy at 64.9% stated confidence, n=116, a 28-point understatement.
The under-confidence failure is a suppressed-signal failure: the gate does not fire often enough because stated confidence lags what the model delivers, and it costs money quietly by declining trades that were good.
The path-passage classifier is a three-class LightGBM returning p_up, p_down and p_none.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quantified but single-author and thinly sampled
The mechanism is transparent and checkable — the EV formulas, the 80/20 time-ordered wrapper fit on 6,988 rows, a seed-42 500-row holdout, per-bin confidence and accuracy figures, and explicit ECE/log-loss deltas. That is well above the norm for a build log. It is nonetheless one self-reported source with no independent replication, the headline over-confidence finding rests on an 11-sample bin, the holdout is drawn from the model's own training data rather than live signals, and the author concedes absolute cross-asset ECE comparisons are informal because BTC's barrier is 150bps versus 200bps elsewhere.
Prototype, not in the live path
The only adoption facts in the cluster are internal to one author's own system, and they are explicitly pre-production: the integration guide status line reads prototype, not integrated into the live inference path; there has been no live-data refit because signals_history does not yet carry realized 24-hour outcomes; and the rollout mechanism (an allow-list env var plus loader fallback) is described as a plan for six of seven assets rather than a shipped state. Migration 028 landing to collect outcomes is real but preparatory. No third-party or organizational uptake is claimed.
Mildly overstated headline, unusually candid body
The title and dek generalise from thin evidence — 'one over-confident asset, six under-confident, no single fix' leans on an 11-sample ETH bin and on a claim that under-confidence 'costs money quietly' for which no realized P&L is shown. That is a modest overstatement. It is largely offset by an unusually honest body: the author flags the prototype status, the absence of live-data refit, the training-data holdout, the informal cross-asset ECE comparison, the LINK failure, and two unreconciled explanations for it. Net effect is a small positive gap rather than a promotional one.
Self-reported author showcase, no commercial ask
Every number is produced and reported by the author about the author's own trading system, published on a developer platform where technical-credibility building is the payoff — a real self-report bias, and there is no independent audit trail. Countervailing factors keep this moderate: no product, vendor, pricing, funding, or licensing interest is present in the cluster, the piece publishes a failing asset and an unresolved diagnosis, and it discloses methodological caveats that a purely promotional post would omit.
Moderate — mechanism solid, generalisation unproven
Confidence in the mechanical core is fairly high: the linearity argument connecting probability error to position sizing follows from the published formulas, and the direction of the per-asset calibration errors is documented with bin-level numbers. Confidence in the broader conclusions is much lower — one publisher, no replication, an 11-sample headline bin, a holdout carved from training data rather than live signals, no realized economic effect, and an unresolved LINK case. The narrow evidence base and pre-production status cap the overall figure below the midpoint.
build
Three manual interventions in a month, and every guard was working as designed1 distinct publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 distinct publisher
build
Force the tool call, then hand Lightsail a long-lived key1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 23, 2026