Build1 distinct publisher3 min readUpdated
FINAL-Bench moved a fixed LightGBM baseline by 0.211 AUROC purely by changing how hERG data was divided. That is roughly eight times the spread across random seeds.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
AUROC starts at 0.5, not at zero. Measured against that floor, the time-split baseline held 0.106 of usable ranking signal and the random-split version of the same model held 0.318, three times as much apparent skill [5][6][1]. The LightGBM defaults, the Morgan fingerprints and the hyperparameters were identical in both runs [8].
Size that against the variation a modeller would normally report. Five random seeds spanned 0.803 to 0.830, a band of 0.027 [6]. The choice of split moved the score about eight times further than reseeding did [7][2]. On a board that does not print its split type, the largest single term in the headline number is the one nobody documented.
The mechanism is in how the data gets generated. A promising scaffold is modified into a family of close analogues, and a random split can distribute that family across training and test, so the model is graded on compounds that resemble examples it has already seen [9]. A chronological split asks it to predict measurements reported after its snapshot of knowledge ends [9]. Both produce an AUROC on hERG. Only one of them is a forecast.
The label side is the second piece of arithmetic. Repeat measurement is common here: of 11,972 records retained after unit checks and parsing, 9,788 were distinct compounds, so roughly 18 percent were re-measurements [4][7]. FINAL-Bench compared hERG values for the same compound across different publications and found, over 5,185 pairs covering 857 compounds, a median absolute difference of 0.148 log units and a mean of 0.475 [12]. A mean sitting three times above the median is the signature of a long tail [4], and the tail is where the damage is: at the 90th percentile two papers disagree by 1.338 log units, more than twenty-fold in potency [13]. From that spread the project estimates a single-measurement standard deviation of 0.421 log units and prints it beside model scores [14].
That floor is why classification got demoted. 696 of the 1,338 test compounds sit within one noise floor of the pIC50 = 5.0 cutoff, and LEADBOARD now scores regression first and classification second [15]. Fifty-two percent of the classification test set [3] could change class if the assay were run again in another lab, which means a ranking on that metric is partly a ranking of measurement luck.
Two limits on the finding. FINAL-Bench does not claim that published drug benchmarks generally use random splits, and notes that scaffold splits are common and set a different level of difficulty [10]; the 0.211 figure comes from one endpoint, built from 12,021 ChEMBL 37 records reduced to 9,788 unique compounds [4][7]. And the evidence is self-published on Hugging Face by the party that runs the board [2], whose founder's credentials, 12 patent applications and four co-authored papers, are sourced to VIDRAFT's own site [16]. The board's authority rests on 18,382 held-out compounds whose labels participants never see [3], which leaves the organiser as the only party able to score it.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
FINAL-Bench extracted 12,021 hERG records with document years from ChEMBL 37, retained 11,972 after unit checks and parsing, and reduced those to 9,788 unique compounds.
A default LightGBM model on Morgan fingerprints, trained on compounds reported before 2022 and tested on compounds reported later, scored 0.606 AUROC.
On a random split of the same dataset the model averaged 0.818 AUROC across five seeds, with individual results ranging from 0.803 to 0.830.
The gap between the time-split and random-split results was 0.211 AUROC.
The model, fingerprints and hyperparameters stayed fixed between the two runs; only the split changed.
Comparing hERG measurements for the same compounds across different publications, across 5,185 pairs covering 857 compounds, the median absolute difference was 0.148 log units and the mean was 0.475.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but single-origin and self-reported
The quantitative core is unusually specific and internally consistent - record counts through the ChEMBL 37 curation pipeline, five-seed ranges, MAE against a constant baseline, 5,185 measurement pairs, a derived 0.421 log-unit noise floor - and the reporting flags which figures come from the project. But every number originates in FINAL-Bench's own August 22 Hugging Face article relayed by one publisher, with no independent replication, external audit or third-party inspection of the withheld labels, so the evidence is rich in detail and thin in corroboration.
Published, pre-submission, no external users named
Adoption evidence is limited to the launch itself: the leaderboard and article are live on Hugging Face with 21 boards and three project-authored baselines published 'before submissions open'. The supplied material names no external participants, labs, companies or institutional users, and no submission counts or downloads, so measured adoption is essentially the release plus the maintainers' own baseline runs.
Mildly overstated by single-origin framing
The substantive claims are stated narrowly and self-limited - the project explicitly declines to say drug benchmarks universally use random splits, notes scaffold splits are common, and publishes a constant baseline that beats its own model on seven of nineteen regression boards, which is anti-hype behavior. The modest positive gap comes from packaging rather than arithmetic: a striking headline delta, an institution-building ambition and self-attested founder credentials are carried on one project-authored artifact with no external submissions or independent replication yet.
Vendor-run ruler with commercial and IP ambitions
The benchmark's author is also a commercial actor: LEADBOARD is a FINAL-Bench/VIDRAFT project, the announcement comes from VIDRAFT's CEO, the founder biography and credential counts come from VIDRAFT's own website, and the stated aim is commercial products and intellectual property to finance a research institution. The coverage notes the project is 'trying to standardize the ruler before building more systems that claim to beat it', which is a direct conflict to price in - partly offset by design choices that make the ruler harder for its own author to game, such as excluding random-split boards, holding labels off-platform and SHA-stamping evaluations.
Moderate: precise numbers, one origin
Confidence is capped by structure, not by internal quality. One publisher and one project-authored primary artifact underpin every figure; the mechanism (analogue leakage under random splits, assay label noise) is well established in the field and the arithmetic checks out, which supports the direction of the finding. What cannot be confirmed from the supplied material is the reproducibility of the specific numbers, the state of the other 20 boards, and any external uptake.
build
The judge is the product: building the scorer for an AI malaria drug leaderboard1 distinct publisher
build
Shanghai AI Lab's 397B science agent shipped in July; the paper explaining it landed August 131 distinct publisher
build
Unsloth's 10% quant claim is really about which machines can run a 27B model1 distinct publisher
build
Your vLLM Manifest Would Boot SGLang Too, And That Is the Problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026