Build1 distinct publisher3 min readUpdated
VIDRAFT and FINAL-Bench opened a public leaderboard for AI-designed PfDHODH inhibitors. The instructive part is the fourteen defects they found in their own scoring system first.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
VIDRAFT and FINAL-Bench have opened a public leaderboard, the Open Discovery Challenge, for AI-designed malaria drug candidates against PfDHODH, the parasite's dihydroorotate dehydrogenase [1][2]. The useful output of the exercise so far is not a molecule; it is a list of fourteen defects the team reports finding while validating the scorer itself [9].
The framing is worth stating plainly, because the marketing rarely does: a generative model can emit valid-looking SMILES indefinitely, and structural plausibility answers none of the questions that decide whether a candidate deserves lab time [4]. The challenge turns six of those questions into published axes: whole-cell kill, binding to PfDHODH, avoidance of the homologous human enzyme, an acceptable ADMET profile, genuine novelty rather than a decorated known scaffold, and realistic synthesizability [5][6]. A viable candidate also has to survive long enough to matter and cross both the red-blood-cell membrane and the parasite membrane [3].
Publishing a rubric is cheap. The reported failures show what it costs to make one behave.
The first toxicity thresholds looked conventional and rejected all three approved antimalarials, plus caffeine [7]. The team traced this to predictors biased against large, lipophilic molecules whose outputs had been converted straight into hard cutoffs [8]. The resulting rule is the most portable thing in the writeup: every threshold must let approved drugs through before it is allowed to reject anyone [10]. Applying it surfaced a molecular-weight cap that excluded a 531.9 Da reference drug and a reactivity detector that kept rejecting an approved compound [11].
Normalization behaved the same way. Raw binding scores reward heavier molecules, so the team normalized by heavy-atom count, a standard correction that produced the opposite bias and let caffeine nearly tie the clinical candidate [12][13]. Adding a potency floor alongside the efficiency term closed it [14]. Transformed metrics do not remove incentives, they move them, which is why each one needs adversarial controls [15].
The activity model is the cleanest case of a data defect wearing a model costume. It predicted caffeine as active at 1 micromolar [16]. The training set kept compounds with measured activity and dropped records such as "no effect at 100 micromolar" because they carried no numeric value, so the model had only ever seen things that worked [17]. After 5,190 failure records were added and the model retrained, separation between the clinical candidate and caffeine widened from 1.00 to 1.74 log units [18], a gain of 0.74 log units, or roughly five and a half fold in predicted potency ratio [19].
Two defects were plumbing. An approved drug with a literature IC50 of 13 nM came back at 248 micromolar, four orders of magnitude off, and the team's first conclusion was that the docking tool could not support absolute values [20]. That conclusion was withdrawn when the same configuration, called directly, returned 13 nM: the integration path was wrong, not the science [21]. Separately, the novelty scorer called the wrong function when reconstructing fingerprints, raised no exception, and returned plausible numbers, so only the novelty axis would have been quietly wrong [22]. A store-read-compare round-trip identity check caught it [23].
Calibration is where the account thins out. Candidates are graded on uncertainty-aware lower bounds rather than point predictions, and a nominally 90% bound measured 83% coverage [24] - a seven-point shortfall in the direction that flatters entrants [25]. The published text breaks off mid-sentence on what the statistical correction assumed [26].
Watch for the six defects not described in the public writeup [27], and for the coverage fix. And watch what happens when entrants start optimizing against the rubric rather than the disease, because that is the real load test for any judge with a leaderboard attached.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The Open Discovery Challenge is a public leaderboard opened by VIDRAFT and FINAL-Bench for AI-designed malaria drug candidates.
The target is PfDHODH, the malaria parasite's dihydroorotate dehydrogenase.
A useful candidate must inhibit the parasite enzyme, avoid the homologous human enzyme, survive long enough to matter, and cross both the red-blood-cell membrane and the parasite membrane.
The listed verification questions are: does it kill the parasite in a whole cell, does it bind PfDHODH strongly enough, does it avoid human DHODH, is its ADMET profile acceptable, is it genuinely novel rather than a decorated known scaffold, and can it realistically be synthesized.
The Open Discovery Challenge turns those questions into a published six-axis scoring system.
A language or molecular model can emit valid-looking SMILES strings indefinitely, but structural plausibility alone does not answer the questions that determine whether a candidate deserves further work; molecule generation has become accessible while verification has not.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but wholly first-party and computational
The writeup is unusually specific — named defects, exact figures (531.9 Da, 1.00 to 1.74 log units, 5,190 failure records, 13 nM versus 248 µM, 83% versus 90.01% coverage) and reproducible validation rules — which is stronger than typical launch material. But every figure comes from one self-published post by the organizers themselves, with no independent audit, no second publisher, and no experimental confirmation that in-silico rankings correspond to real potency or selectivity. Six of the fourteen reported defects are never described at all.
Launched and publicly documented, uptake unmeasured
There is a concrete, dated public artifact: an open leaderboard with a published rubric, announcement and entrant guide on Hugging Face, plus a documented internal validation run against approved drugs and inert controls. That is more than an announcement of intent. But the cluster contains no submission counts, participant names, external deployments, or third-party use of the scorer, so adoption beyond the organizers cannot be scored higher.
Modestly overstated by framing, restrained in substance
The post is markedly self-critical — it publishes its own fourteen defects, withdraws an earlier wrong conclusion about the docking tool, and reports measured rather than nominal coverage, all of which pull the gap toward zero or negative. The residual overstatement is in framing: a 'verifiable judge' for drug candidates is built entirely from predictive models whose own biases the article documents, and nothing here shows that a high leaderboard score corresponds to real-world antimalarial activity. Ranking a clinical candidate above inert controls is a sanity check, not validation of the scorer's absolute claims.
Organizer-authored recruitment content
The only source is written from the organizers' side of the Open Discovery Challenge and functions in part as an entrant funnel: it links the announcement and entrant guide, stresses that no proprietary molecular model is required, and explains the rubric contestants will be scored against. That is a clear interest in attracting submissions and in establishing the scorer as authoritative. The candid defect disclosure and the availability of a published rubric and positive controls partially offset the incentive, since both create externally checkable commitments.
Coherent single-source account, uncorroborated
Internal coherence is high and the technical narrative is specific enough to be falsifiable, which supports moderate confidence in the engineering claims as descriptions of what the team did. Confidence is capped by having one publisher, one author-side viewpoint, zero independent replication, a partially truncated body, and a ledger claim about the text's ending that the supplied source contradicts.
build
Unsloth's 10% quant claim is really about which machines can run a 27B model1 distinct publisher
build
Your vLLM Manifest Would Boot SGLang Too, And That Is the Problem1 distinct publisher
build
Vivodyne is spending venture money on wet-lab throughput, not bigger models2 distinct publishers
build
The click succeeded and nothing happened: your agent harness needs an injected canary1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 14, 2026