Build1 publisher3 min readPublished Updated
The judge is the product: building the scorer for an AI malaria drug leaderboard
VIDRAFT and FINAL-Bench opened a public leaderboard for AI-designed PfDHODH inhibitors. The instructive part is the fourteen defects they found in their own scoring system first.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The Open Discovery Challenge is a public leaderboard opened by VIDRAFT and FINAL-Bench for AI-designed malaria drug candidates.
- The target is PfDHODH, the malaria parasite's dihydroorotate dehydrogenase.
- A useful candidate must inhibit the parasite enzyme, avoid the homologous human enzyme, survive long enough to matter, and cross both the red-blood-cell membrane and the parasite membrane.
- A language or molecular model can emit valid-looking SMILES strings indefinitely, but structural plausibility alone does not answer the questions that determine whether a candidate deserves further work; molecule generation has become accessible while verification has not.
- The listed verification questions are: does it kill the parasite in a whole cell, does it bind PfDHODH strongly enough, does it avoid human DHODH, is its ADMET profile acceptable, is it genuinely novel rather than a decorated known scaffold, and can it realistically be synthesized.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
VIDRAFT and FINAL-Bench have opened a public leaderboard, the Open Discovery Challenge, for AI-designed malaria drug candidates against PfDHODH, the parasite's dihydroorotate dehydrogenase [1][2]. The useful output of the exercise so far is not a molecule; it is a list of fourteen defects the team reports finding while validating the scorer itself [9].
The framing is worth stating plainly, because the marketing rarely does: a generative model can emit valid-looking SMILES indefinitely, and structural plausibility answers none of the questions that decide whether a candidate deserves lab time [4]. The challenge turns six of those questions into published axes: whole-cell kill, binding to PfDHODH, avoidance of the homologous human enzyme, an acceptable ADMET profile, genuine novelty rather than a decorated known scaffold, and realistic synthesizability [5][6]. A viable candidate also has to survive long enough to matter and cross both the red-blood-cell membrane and the parasite membrane [3].
Publishing a rubric is cheap. The reported failures show what it costs to make one behave.
The first toxicity thresholds looked conventional and rejected all three approved antimalarials, plus caffeine [7]. The team traced this to predictors biased against large, lipophilic molecules whose outputs had been converted straight into hard cutoffs [8]. The resulting rule is the most portable thing in the writeup: every threshold must let approved drugs through before it is allowed to reject anyone [10]. Applying it surfaced a molecular-weight cap that excluded a 531.9 Da reference drug and a reactivity detector that kept rejecting an approved compound [11].
Normalization behaved the same way. Raw binding scores reward heavier molecules, so the team normalized by heavy-atom count, a standard correction that produced the opposite bias and let caffeine nearly tie the clinical candidate [12][13]. Adding a potency floor alongside the efficiency term closed it [14]. Transformed metrics do not remove incentives, they move them, which is why each one needs adversarial controls [15].
The activity model is the cleanest case of a data defect wearing a model costume. It predicted caffeine as active at 1 micromolar [16]. The training set kept compounds with measured activity and dropped records such as "no effect at 100 micromolar" because they carried no numeric value, so the model had only ever seen things that worked [17]. After 5,190 failure records were added and the model retrained, separation between the clinical candidate and caffeine widened from 1.00 to 1.74 log units [18], a gain of 0.74 log units, or roughly five and a half fold in predicted potency ratio [19].
Two defects were plumbing. An approved drug with a literature IC50 of 13 nM came back at 248 micromolar, four orders of magnitude off, and the team's first conclusion was that the docking tool could not support absolute values [20]. That conclusion was withdrawn when the same configuration, called directly, returned 13 nM: the integration path was wrong, not the science [21]. Separately, the novelty scorer called the wrong function when reconstructing fingerprints, raised no exception, and returned plausible numbers, so only the novelty axis would have been quietly wrong [22]. A store-read-compare round-trip identity check caught it [23].
Calibration is where the account thins out. Candidates are graded on uncertainty-aware lower bounds rather than point predictions, and a nominally 90% bound measured 83% coverage [24] - a seven-point shortfall in the direction that flatters entrants [25]. The published text breaks off mid-sentence on what the statistical correction assumed [26].
Watch for the six defects not described in the public writeup [27], and for the coverage fix. And watch what happens when entrants start optimizing against the rubric rather than the disease, because that is the real load test for any judge with a leaderboard attached.