Skip to content

Build1 publisher2 min readPublished

Netflix scored a new launch rule by replaying 123 finished A/B tests

A new paper evaluates the decision rule itself by asking what every past test would have returned under it, and reports that the obvious backward-looking estimate runs high whenever each experiment's signal is weak.

The Engineer · Build desk

Illustration accompanying Netflix scored a new launch rule by replaying 123 finished A/B tests

What happened

  • A new paper proposes judging an experimentation program's decision rule by its cumulative return to a north star metric, a quantity no single test measures but a corpus of past tests can estimate.
  • The straightforward backward-looking tally is biased by a winner's curse, because arms that won partly on noise have true effects smaller than their recorded lifts.
  • In a case study of 123 historical Netflix A/B tests, the method estimated a higher cumulative return for a new decision rule, and the rule was adopted.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Scoring rules this way needs the north star measured per arm in finished tests, which is the same measurement problem proxy metrics were adopted to work around.
  • decision The object under review becomes which threshold, proxy and guardrail combination the organisation standardises on, and that combination can be compared before it ships.
  • exposure A team that backtests candidate rules with the plug-in estimator will tend to select whichever rule best exploited past noise, and pay for it on the next thousand tests.
  • precedent At Netflix an estimate of cumulative returns was sufficient evidence to change a company-wide launch procedure, giving other programs a reference for what a rule change has to show.

A decision rule picks a treatment arm using a measured lift [3]. When the noise in each test is large next to the true effects, some arms get picked because the noise happened to favour them [5]. Replay the history under that rule and the tally credits it with those noisy lifts, so the backward-looking number runs high [5]. The paper reports that this bias survives an infinite number of experiments as long as each experiment has a finite sample size, and gets worse as signal-to-noise falls [6]. Adding more tests to the corpus does not remove it [6].

The thing being scored is a whole standard operating procedure: statistical significance thresholds, proxy metrics blessed because they carry more signal, guardrail metrics that block a launch when something like app crashes or customer service contacts moves the wrong way, and surrogate indices that collapse several proxies into one scalar [11]. Each of those is a knob someone can tune against history. The authors say rigorous guidance on how to evaluate and choose decision rules is scarce [13], and warn that a rule exploiting spurious correlations can look strong on past experiments and generalise poorly to future ones [7].

Their remedy is described in the abstract this way: "We develop a cross-validation estimator that is much less biased than the naive plug-in estimator under conditions realistic to digital experimentation" [8]. The case study covers 123 historical A/B tests at Netflix [9]. That corpus is smaller than a single year of Netflix's own program, which the paper puts at thousands of tests per year [14].

For this to run on someone else's program, the corpus has to exist as data, with per-arm outcomes still retrievable from finished tests [4]. The north star also has to be measurable in those tests, and that is the awkward part: proxy metrics get adopted in the first place because they have a higher signal-to-noise ratio than the north star, or are easier to measure in short, sample-constrained experiments [12]. A program that logged only the proxy and the launch decision has nothing to compute cumulative returns against.

The arxiv HTML renders the headline effect size as an empty space, which is an unlucky number to lose in typesetting [10]. The same sentence says the estimate led directly to the adoption of the new rule [9].

What to watch

  • Whether a later version of the paper renders the estimated increase in cumulative returns, and how large it is.
  • Whether the cross-validation estimator holds up on corpora much smaller than 123 experiments, where most programs sit.
  • Whether Netflix publishes the composition of the adopted rule: the threshold, the proxy metrics and the guardrails it uses.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories