Build1 publisher3 min readPublished
The bot returned 4.33%. Doing nothing returned 127.77%. The useful part is why.
A developer published the losing number instead of the equity curve, along with the four specific mechanisms that had made an unprofitable strategy look profitable.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author's best strategy returned +4.33% over three years of out-of-sample testing, while buying Bitcoin once and holding it returned +127.77%.
- The gap between buy-and-hold and the strategy is 123.44 percentage points.
- The author states the original backtester said the strategy was profitable and was wrong in four separate ways, all endemic in hobby trading bots, and rebuilt the system so that lying was impossible.
- The system, called TradingAI, is about 4,000 lines of Python: a strategy engine, a backtester, a walk-forward validation harness, a live paper-trading bot with Telegram alerts, a real-time dashboard, and 130+ tests running in CI.
- The traded rule is a Donchian channel breakout: buy 20-day highs, exit on a trailing stop, using real exchange data via ccxt.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer writing on dev.to published a three-year out-of-sample result for a crypto trading system: +4.33%, against +127.77% for buying Bitcoin once and holding it [1]. That is a gap of 123.44 percentage points [2], and the post is worth reading because most of its length goes to the four ways the same strategy looked profitable before the test harness was rebuilt [3].
The thing under test is not exotic. About 4,000 lines of Python: a strategy engine, a backtester, a walk-forward validation harness, a paper-trading bot with Telegram alerts, a dashboard, and 130-plus tests in CI [4]. The rule is a Donchian channel breakout, buy 20-day highs and exit on a trailing stop, run on exchange data through ccxt [5].
Lie one was look-ahead. The old code computed a signal from a candle and then bought at that same candle's close [6]. The fix was structural rather than a patch: one engine enforces the ordering for every strategy, and a test feeds in a strategy that deliberately peeks one bar ahead and asserts it cannot profit [7]. That is the part any team can copy tomorrow. The bug becomes a permanent adversarial test, so the next contributor cannot reintroduce it quietly.
Lie two was fees, which the original backtester set to zero [8]. Kraken's taker fee is 0.26% per side [9], or 0.52% for a round trip [10]. Re-run with honest costs, the same signal produced: on daily bars, 14 trades, 1.3% of account per year in fees, +4.33%; on 4-hour bars, 61 trades, 11.4% in fees, -6.01%; on 1-minute bars, 52 trades in 45 days, 23.5% in fees over those 45 days, -25.18% and not a single winning trade [11]. The signal was identical at every speed [12]. On 1-minute bars the author puts the fee at roughly six times the average move being targeted [13]. Even the surviving configuration is thin: 1.3% a year in fees against roughly 1.44% a year net implies fees ate about 47% of gross return [14]. One thing the post leaves unexplained is that the 4-hour run has about 4.4 times the trades of the daily run but about 8.8 times the annual fee drag [15], which anyone reproducing this should reconcile against window length and position sizing.
Lie three was tuning and testing on the same data. The +4.33% comes from walk-forward analysis: optimise on a training window, trade the next window the optimiser never saw, roll forward, and score only unseen days across five years of data [16].
Lie four was luck. A permutation test shuffled daily returns 200 times and re-ran the strategy on each scrambled history; the strategy beat 97.5% of the shuffles, p = 0.025 [17]. The author's own reading: the edge is real, small, and smaller than the fees [18].
Three priors died on the same bench. Seven summed indicators finished last against a one-rule strategy, -2.4% versus +3.0% [19]. Equal sleeves of BTC, ETH and SOL produced a portfolio Sharpe of 0.30 against 0.38 for BTC alone [20]; the author kept the feature anyway, on the grounds that a bench that spares its author's ideas is worthless [21]. The one defensible win: through a stretch where Bitcoin drew down 53%, the bot's worst drawdown was 4% [22].
Watch whether the anti-leak test stays green as strategies accumulate, and whether the drawdown behaviour survives a live down leg rather than a scored one.