Skip to content

Science1 publisher3 min readPublished

Rice statisticians' Bayesian model attaches uncertainty estimates to merged baseball rankings

Rice and Cornell statisticians merged four years of top-100 lists from five outlets into an MLB consensus ranking that reports how certain each placement is. The design carries over to any team pooling partial, disagreeing rankings, though its real-world test covered just 55 baseball players.

The Scientist · Science desk

Illustration accompanying Rice statisticians' Bayesian model attaches uncertainty estimates to merged baseball rankings

What happened

  • The model treats each list as evidence about a hidden player value, accepts partial lists and ties, borrows information across years, and scores how closely each outlet tracks the consensus.
  • Wins above replacement had the strongest positive association with rank, age was close behind with younger players favored, and higher base salary also went with higher rank.
  • The model's 2023 top four were Shohei Ohtani, Aaron Judge, Mike Trout and Mookie Betts, and all four were selected as All-Star starters that season.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • capability A pooled ranking built this way can flag pairs the data cannot separate, so a choice between two near-equal candidates need not rest on an order the evidence does not support.
  • constraint Because the test pool was restricted to consistently listed players, the model's behavior on sparse, high-churn lists, where merging is hardest, remains unshown.
  • decision A team adopting BMRR outside baseball will need its own outcome measure to validate the consensus, since the reported check is agreement with All-Star selections.

Averaging is the obvious way to merge lists that disagree. According to the Rice team, it hides information when a player is left off one list entirely or when an outlet's view of him shifts sharply between seasons [5]. BMRR instead treats every ranking as evidence about an unobserved value for each player and pools that evidence in a hierarchical Bayesian model [6]. The raw material here was five outlets over four years: 20 lists and 2,000 ranked slots [3][1].

"Rankings look simple on the surface, but statistically they're actually very complicated," said Rose Graves, a statistics doctoral student at Rice and corresponding author of the paper in the Journal of Quantitative Analysis of Sports [8][1]. She added: "Two experts may rank different numbers of items, disagree about the order or even change their opinions over time. We wanted to create a way to bring all of that information together while accounting for uncertainty." [9]

The output is a ranking with an uncertainty measure attached. It shows where the data strongly favor one player and where two are too close to separate [7]. The model also estimates how closely each outlet tracks the consensus [6]. "Rather than asking only who ranks first, this framework lets us ask how confident we are in that conclusion, why someone may be ranked highly and how much agreement actually exists among the people doing the ranking," said Marina Vannucci, the Noah Harding Professor of Statistics at Rice and a co-author [10].

The "why" came from three player traits fed into the model: age, base salary and wins above replacement, or WAR, a statistic estimating a player's overall contribution to his team [13]. WAR had the strongest positive association with rank [14]. Age was close behind in magnitude, with younger players ranked more favorably, and higher salary also went with higher rank [14]. The authors take this as a sign that media rankings reflect expectations about future potential and value as well as recent performance [15]. Salary is harder to read. The same account notes that perceptions of player value influence salary negotiations and arbitration [12], so a high ranking and a high salary can each feed the other.

For 2023, the model's four highest-ranked players were Shohei Ohtani, Aaron Judge, Mike Trout and Mookie Betts, and all four were selected as All-Star starters that season [16]. Agreement with another human selection is a sensible check on a model built from human opinion. The thing this doesn't tell you is whether the BMRR consensus predicts anything, next season's WAR for instance, better than a plain average of the same 20 lists would [1]. The published account reports no such comparison.

In my view the reusable part is the model's structure. Partial lists, ties, sources that agree with the consensus to different degrees, and repeated rounds of judgment are what BMRR was built to absorb [2][6]. Those features show up wherever several people each rank their own shortlist. "That kind of uncertainty quantification can be extremely important when rankings are being used to inform decisions," Vannucci said [11].

The condition is the test pool. It held only the 55 players who made at least one outlet's list in every one of the four years [4]. Players who broke in or dropped off mid-period, the incomplete-list case that makes averaging fail [5], were left out of it [4].

What to watch

  • A test of whether the BMRR consensus predicts next-season performance better than a plain average of the same lists.
  • An application of BMRR to sparse, high-churn rankings where few items appear on every list, outside sports.
  • Whether the Rice and Cornell authors release code that lets other groups run the model on their own ranking data.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories