"Explained label variation" would have to be a variance decomposition: fit the labels against a set of factors, set aside the part nothing accounts for, then report how much of the remainder one factor owns. That construction hides its own denominator. A prompt sentence can own most of the explained share while the explained share itself is thin, and the two readings imply opposite things about how much eval wording costs a buyer. Nothing in the supplied text sets out the label set, the models compared, or the two prompt variants, so there is no way to tell which reading applies [10].
The awkward part is that the essay predicts this behaviour anyway. Its authors argue there is no shared, absolute definition of ground truth in medicine, and that any eval built on clinical data ends up scored against what one physician did on one particular day [4]. If that holds, a benchmark number is partly a property of the sentence used to elicit the label, which is close to what a16z and Protege say outright: a good benchmark result is as much a statement about how you wrote the test as about clinical merit against a different model or a different evaluation [3]. A ranking that reverses on a rewrite is the described system working as described.
What the text does put on the record is smaller and checkable. The sampling was millions of records drawn from nearly a trillion tokens of EMR text across billions of notes, labelled descriptive rather than causal [5]. Part way through 2026, about 1,686 notes per million carried a patient asking for clarification on something a model had said [6], roughly one note in 593 [11]. That sentence is where both copies run out [8][9]. The remedy the authors endorse is the slow one, following the patient forward long enough to see who was right [13], with Protege positioned to do it [14].