Skip to content

Build1 publisher3 min readPublished

The best model on CommentBench rediscovered 8.3% of the points human reviewers raised

CommentBench splits human comments on AI-safety posts and drafts into target points, filters out the ones a model could not reach without extra context, and has Opus 5 judge which of the rest a model hit.

The Engineer · Build desk

Photograph accompanying The best model on CommentBench rediscovered 8.3% of the points human reviewers raised
Photo: lesswrong.com

What happened

  • Fable 5 topped CommentBench, Fable 5.1 followed at 7.5%, and the authors report Astra at 5.5% and Sol at 5.3% as an underperformance for OpenAI models against their other capability benchmarks.
  • Model scores line up across settings at Pearson r = 0.86 to 0.97, covering forum posts, shortforms, research drafts and replies.
  • The team says its current plan is to share the benchmark selectively with AI labs.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Credit flows only for rediscovering a point the humans made, so a model comment the authors themselves call insightful adds nothing to the score.
  • decision Anyone deciding whether a model can stand in for a senior reviewer has to pay for the human arm of the experiment, because the post puts that measurement in the future tense.
  • contradiction The authors' stated view of what frontier models can already do with elicitation sits well above the highest number their own pipeline produced.
  • precedent A benchmark handed to labs privately can be trained against while outsiders have no way to reproduce the 8.3%.

The score is recall on a curated denominator. For each document, CommentBench computes the share of human target points matched by at least one model-written comment, and each reported line is the mean of that share across documents, averaged over four samples [7]. Before the models see anything, every human comment is split into target points, and the points a model could not make without additional context are filtered out [5]. So the denominator is the subset of human feedback a model had a fair shot at. Filtering that way raises the score; unfiltered human comments would give a lower one [5].

Matching is judged by a model. Opus 5 is the "matcher" that decides how many human target points the AI comments hit [6]. One model's comments get graded by another model's reading of what the humans meant. Nothing in the setup charges a model for a comment that is wrong, and nothing credits one that is new. Some models produce insightful comments, including ones the humans did not make, according to the post [12]. Those score zero.

Take 8.3% at face value and roughly 92% of the feasible human points went unmatched by the strongest model in the set [1][17]. Fable 5's rate is about 1.5 times Astra's 5.5% [18]. Both figures pool four samples per document, with credit for any single comment landing a point, so they describe several attempts together and not one review pass [7].

Reading 8.3% as a ceiling on model review needs a human number, and the post does not report one. According to the post, benchmarking model performance against humans would mean getting human researchers to leave comments under comparable conditions, and it describes commenting on research drafts as a fuzzy task often performed by senior researchers [15][19]. Until a second senior reviewer goes through the same pipeline, the gap is unmeasured. A human graded on rediscovering another human's points might also land in single digits.

The figure transfers to your review process only if three things hold. Your documents have to resemble the corpus: LessWrong and EA Forum posts, forum shortforms and private Google Doc research drafts selected for high-quality comments [4]. Your definition of a useful review has to be overlap with what those particular commenters wrote. And Opus 5's judgement of a match has to agree with yours [6].

Some of the design is solid. Scores correlate across pairs of settings at Pearson r = 0.86 to 0.97, so the ranking of models holds whether the document is a forum post, a shortform, a draft or a reply [9]. The methodology was calibrated on the Google Docs setting and generalised to the others with minimal adaptation [11]. The team also checked memorisation, found no consistent advantage on posts published before a model's knowledge cutoff, and notes that every public document postdates Fable 5's cutoff [10].

The authors rate current models above what the benchmark measures. "With some elicitation, we think frontier models are already good enough to provide useful conceptual feedback on drafts," they wrote [13]. CommentBench is the validation step for that claim, and the current plan is to share it selectively with AI labs [14]. They expect hill-climbing CommentBench to select for making AI comments more human-like [16].

What to watch

  • Whether the authors run senior human reviewers through the same target-matching pipeline to produce a human baseline score.
  • Which labs receive the benchmark, and whether reported match rates move after they do.
  • Whether a precision measure appears that scores model comments the humans never made.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories