Skip to content

Science1 publisher2 min readPublished

Simulated skewed data favours the BCa bootstrap over the textbook confidence interval

The author of free range statistics resampled heavily skewed distributions and reports that a bias-corrected bootstrap holds more of its promised 95% coverage than the interval taught in intro courses, with both still short of 95%.

The Scientist · Science desk

Illustration accompanying Simulated skewed data favours the BCa bootstrap over the textbook confidence interval

What happened

  • The follow-up simulation compared 95% confidence intervals for the mean built the way basic statistics courses teach with intervals from the bias-corrected and adjusted bootstrap, on a few heavily skewed distributions.
  • The author reports that the bootstrap did considerably better than relying on the central limit theorem, with the clearest gain at the smaller sample sizes.
  • Both methods still contained the true mean a lot less often than 95% of the time, the level each interval was constructed to hit.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint The better method still under-covers on the same data, so switching interval formulas does not restore a 95% claim on a heavily skewed metric; anyone who wants honest uncertainty has to widen the interval or estimate something other than the mean.
  • exposure Analysts who treat a thousand observations as safely large enough for the central limit theorem are quoting intervals that miss the true mean more often than one time in twenty, and the direction of that error understates risk.
  • decision The switch to BCa pays most where samples are small, so the gain is bigger for someone reporting on dozens of observations than for someone reporting on tens of thousands.
  • capability Because the simulation code is published, you can substitute your own distribution and sample size and measure your own coverage instead of assuming the textbook interval holds.

Coverage is the property being tested: over many simulated samples, how often does a 95% interval actually contain the true mean. The author of free range statistics picked the resampling method he expected to do best in this setting, and said why. "No surprise here; from what I understand of the history, this is pretty much exactly what the BCa bootstrap was developed for," he wrote [7]. He described the simulations as the method's home ground, where it does relatively well [8]. Both intervals still came in under 95% [5].

The post does not put a number on either shortfall. The results appear as a displayed set of results, with no coverage percentages, distribution names or replicate counts in the text [10]. "But those actual coverage numbers are still well below 95%, for both methods," he wrote [6].

A nominal 95% interval is built to miss the true mean in about one sample in twenty [15]. Coverage well below 95% means the interval is too narrow and misses more often than that. How much more often depends on the distribution and the sample size. The error runs in one direction: the reported uncertainty is smaller than the real uncertainty.

The sample sizes in play are wide. He lists 10, 30, 200, 1,000, 10,000 and 50,000, a span of 5,000-fold from smallest to largest [11][16], and writes that at those sizes "there's just limits to what you can do" [11]. His earlier post concluded that for these distributions the sample mean can need tens of thousands of observations before its distribution is normal [1]. On that evidence, n = 1,000 alone does not justify the textbook formula.

One piece is still open. The asymmetry of these intervals was raised by Professor Harrell after the earlier post, and it has not been simulated [12]. "I think I've run out of oomph for looking at that, but it is actually an important point to remember," the author wrote [13].

He recommends resampling anyway: "I thoroughly recommend it. It's good stuff." [14] The study simulated a few heavily skewed distributions [2]. Whether revenue per customer or session length at the sample size you have behaves like the distributions here is a question about that data, and the code that produced these results is published [9].

What to watch

  • A table of the actual coverage figures, from the author or from anyone rerunning the published code, would let readers size the gap at their own sample size.
  • A run that reports the sample size at which the bootstrap's advantage over the textbook interval disappears would tell analysts when the choice stops mattering.
  • Naming the specific skewed distributions simulated would show how close they are to the metrics businesses actually average.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories