Build1 publisher3 min readPublished
Most Bayesian modes ship a flat prior stopped on a posterior-probability threshold, a pairing Spotify's engineers say reproduces the false positive rate of frequentist peeking while sounding easier to interpret.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One self-published account, no vendor check
Every technical claim here comes from the company whose existing choice the conclusion happens to endorse. The martingale argument for Bayes factor stopping is checkable against standard statistics; the sentence doing the most work in the piece is not, because Spotify's characterisation of what eight competing platforms ship as their default is nowhere matched to those platforms' own documentation. The paper the post condenses is named but not shown, and the text we hold cuts off partway through the Bayes factor section.
Wide distribution, unmeasured usage
Bayesian modes are shipping across eight named commercial platforms, which is real distribution, and it is the only adoption fact available. How many teams switch those modes on, which configuration they land on, and how many stop on a posterior threshold all go unreported. The one usage disclosure we do have points the other way: Spotify runs frequentist tooling and says it sees no reason to add Bayes.
Framing runs hotter than the finding
The title announces a refusal while the argument lands on convergence: Spotify's own summary is that the two frameworks sit closer together than the debate implies, and it credits Bayes factor stopping and empirical Bayes priors with guarantees it does not dispute. Where the piece overreaches is quantitative reach, not rhetoric. 'The default in almost every platform offering Bayes' carries the whole practical warning and is never broken down platform by platform.
Interested party, interest in plain view
Spotify maintains frequentist experimentation tooling and has a paper to place, and the post's conclusion is that its current stack needs nothing added. The interest is stated rather than buried, and parts of the argument cut against it by conceding that Bayes factor stopping and calibrated priors deliver guarantees a frequentist program would have to match. What remains unbalanced is that the vendors whose defaults are being characterised get no say.
Firm on the position, thin on the products
What Spotify thinks, and why, is documented in its own words and needs no corroboration. Whether the eight platforms behave as described is a separate question our coverage does not answer, and the statistical reasoning arrives in summary form from a paper we cannot read. That split sets the ceiling.
Set the stopping rule to a Bayes factor threshold and you get false positive control even under repeated looks, because the Bayes factor is a martingale under the null [6]. A calibrated empirical Bayes prior instead shrinks effect estimates to counter winner's curse bias and bounds the false discovery rate [7]. Decision-theoretic formulations derive stopping rules that minimise a chosen cost function, and sometimes control frequentist rates, sometimes not [8]. Four configurations carry four different sets of guarantees. By the post's own account, the one shipped as the default is the only one of the four that reproduces peeking [5][17]. The framework, as Spotify describes it, spans state-of-the-art sequential methods and procedures numerically identical to bad frequentist practice [9].
That is the substance behind the interpretability pitch. The argument Spotify quotes is that Bayes is modern, flexible and easier to interpret, and that frequentist complications like multiple testing and sequential testing fall out naturally in the Bayesian way of thinking [2]. Spotify's answer is that this is a theme of oversimplification that leads to inference no better than peeking [3]. Easier to interpret describes the sentence on the dashboard. The false positive rate is set by the stopping rule, not by that description.
The load-bearing argument is smaller than the Bayes-versus-frequentist framing suggests. False positives and false negatives are a classification of outcomes that any experiment program accumulates over time whatever philosophy its inference rests on, so the question is whether controlling those rates is a goal and whether the chosen configuration has parameters that control them [12]. Spotify puts goals first: limiting the number of shipped features with no effect, estimating impact to a stated precision, or minimising a cost function over time, with the statistical configuration following from those goals and constraints [13]. That ordering, reversed, produces the confusion the post describes, where a fact about one goal paired with one configuration is repeated as a fact about Bayesian A/B testing [15].
The taxonomy carries over; the verdict does not. Spotify's stated conclusion is about its own program being served by tooling it already runs [11]. If your program's goal is a cost function minimised over time, the branch that addresses it is the decision-theoretic one [8][13], and the post also argues those Bayesian formulations can be translated into equivalent frequentist formulations [14]. The live choice is which parameters you can set.
The check is cheap: a platform's Bayesian documentation states two values, the default prior and the stopping rule. Flat prior, posterior-probability threshold, dashboard refreshed every morning, and that is the configuration the post says reproduces the error rate of peeking [5]. A Bayes factor threshold or a calibrated prior is a different procedure with different guarantees, sold under the same word [6][7].
Ranked by verification strength, evidence, and original report placement.
GrowthBook, LaunchDarkly, PostHog, Amplitude Experiment, Optimizely, VWO, Statsig and Eppo all ship a Bayesian mode for A/B testing.
The main argument offered for Bayesian A/B testing is that it is modern, flexible and easier to interpret compared to frequentist statistics, and it is often claimed that frequentist complexities like multiple testing and sequential testing fall out naturally in the Bayesian way of thinking.
Spotify argues there is a theme of oversimplification in Bayesian A/B testing discourse that frequently leads to bad inference practice that is no better than peeking in frequentist statistics.
Bayesian inference for A/B testing is not one thing but a family of configurations, each defined by a stopping rule, a prior and a likelihood.
Bayes factor stopping compares the evidence for a treatment effect against the null; because the Bayes factor is a martingale under the null, Bayes factor thresholding provides false positive rate control under peeking, also called optional stopping.
A calibrated empirical Bayes prior can shrink effect estimates to counter winner's curse bias and bound the false discovery rate.
build
Agent-written PRs move the bottleneck to whoever still has to read them1 publisher
product
ARIA makes human authorship a chart eligibility test, and names the three roles that count1 publisher
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 publisher
invest
Local acts squeezed foreign music to 3 per cent of Indonesia's Spotify top 101 publisher
Publishers with included, body-backed reporting in this cluster.
1 article · September 8, 2026