Skip to content

Science1 publisher3 min readPublished

University of Vienna researchers use Benford's law to rank which datasets need closer scrutiny

Johannes Kirchmair's Vienna team built SBBE, a Bayesian check that scores large datasets against Benford's law and flags ones needing scrutiny. Its authors pitch it as triage, picking which data gets a human look before it feeds drug-discovery AI models.

The Scientist · Science desk

Illustration accompanying University of Vienna researchers use Benford's law to rank which datasets need closer scrutiny

What happened

  • The check rests on Benford's law: in many naturally occurring datasets some digits lead more often than others, and a large departure can signal quality problems.
  • Tested on bioactivity data, which record how strongly compounds interact with their target proteins, the method found large databases generally of high quality.
  • It did pick out certain datasets with unusual statistics, and the researchers traced those features to experimental design, data processing or data curation.
  • The work by Uday Abu-Shehab, Matthias Welsch and Kirchmair, from the university's Christian Doppler Laboratory for Molecular Informatics, is published in Patterns.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • decision Teams building AI training sets from pooled databases get an order for manual review, so reviewer time goes first to the datasets the score marks as most irregular.
  • constraint Flags in the bioactivity test came from design and curation choices, so a high score cannot justify dropping data automatically; each flag still needs someone to find its cause.
  • capability A small curated set and a large database can now be ranked on one scale, a comparison that scores sensitive to sample size made unreliable.

The hard part of any Benford screen is deciding how big a departure from the expected digit pattern has to be before it counts as suspicious. "With many existing approaches, the results of quality analyses depend heavily on the number of available data points," Abu-Shehab said [5]. The same raw deviation means one thing in a small dataset and another in a large one [5].

SBBE deals with this by simulating. The team runs extensive computer simulations and feeds them into Bayesian statistics, so each dataset gets a quality estimate with an uncertainty attached [4]. "Our method explicitly takes this uncertainty into account, thereby enabling fair comparisons between datasets of different sizes," Abu-Shehab said [6]. The method is built to run automatically over large collections [1]. Across a pooled collection, the uncertainty term is meant to stop a small dataset from ranking as irregular on sampling noise alone [6].

Kirchmair is careful about what a high score means. "Our method does not provide direct evidence that data is flawed or falsified," he said [7]. "Rather, it helps to filter out datasets from large collections that should be subjected to closer scrutiny. This enables researchers to focus their attention specifically on potentially problematic datasets." [8] The bioactivity test fits that description. The unusual datasets it surfaced had explanations in how the experiments were designed or how the data were processed and curated [11]. Some of those causes are choices a lab made on purpose. A score can point to where a question sits, but a person still has to find out whether the cause is a mistake or a design choice [1].

The thing this doesn't tell you is the hit rate. The press release does not report how many datasets were screened or flagged, how many flags turned out to be real errors, or whether removing flagged data changed any AI model's results [16]. A team would need those figures to judge the tool against the problem the authors describe, where millions of experimental measurements are pulled from databases and fed into AI models for drug discovery [12]. The full paper is published in Patterns [15].

The authors say SBBE is not tied to one field and could apply wherever large volumes of numerical data are generated [14]. I think one condition belongs with that claim. According to the team, Benford's pattern holds in many naturally occurring datasets [3]. Many is not all, so the screen only tells you something about data that should follow the pattern to begin with [2]. Within that limit, Welsch put the goal simply. "Our aim is to provide researchers with an easy-to-use tool to identify anomalies at an early stage and further improve the quality of scientific data collections," he said [13].

What to watch

  • Whether the Patterns paper or a follow-up reports how often SBBE flags matched genuine errors, ideally measured on datasets with known, deliberately planted faults.
  • Whether removing or correcting SBBE-flagged datasets measurably changes the accuracy of drug-discovery AI models trained on them.
  • Whether bioactivity database curators adopt SBBE as a routine check on newly deposited data.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories