Skip to content

Build1 publisher3 min readPublished

GitHub's ReviewBench grades AI code reviewers against 219 public pull requests

GitHub released ReviewBench, an open benchmark that scores AI code reviewers on 219 pull requests from 187 public repositories in 19 languages. Buyers can rerun any reviewer on the public dataset and published judge, though every score depends on one LLM's view of what counts as relevant.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Claude Sonnet 5 grades every submission and accepts a finding only if it is true, relevant and non-trivial, and GitHub publishes the rubric and the judge.
  • GitHub says the benchmark made its offline tests of Copilot code review better at anticipating the direction of production experiments.
  • The full dataset is public, and the post documents how to onboard another code review system and submit results.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Competing reviewers can be run on the same 219 pull requests under the same published judge, so a bake-off no longer depends on whichever repository a vendor chose for its demo.
  • constraint With about 11.5 PRs per language on average, per-language and per-severity results rest on small samples, and a team working mainly in a less common language gets very few data points.
  • decision One LLM grader decides what counts as relevant and non-trivial, so a buyer must choose between accepting that line and re-grading a sample of findings with its own engineers before trusting a ranking.

The best engineering in ReviewBench is how the answer key gets built. GitHub wrote: "No single reviewer, whether human or model, can identify everything worth finding in a pull request." [11] Candidate findings come from four sources: human reviewers, issues inferred from the author's follow-up commits, deterministic analysis tools, and several frontier LLMs from different model families [5]. Findings that describe the same underlying issue are then merged [6]. That step matters. Without it, a bug every model can spot would be counted several times in the golden set, and a reviewer that thinks like the consensus would look stronger than it is.

Scoring is where I would spend review time. One model, Claude Sonnet 5, grades every submission, and a finding counts only if it is true, relevant and non-trivial [7]. The metrics are built to measure newly discovered issues as well as known ones [9]. That implies the grader also rules on findings outside the golden set, with no human label to check against. Relevant and non-trivial are judgment calls. A team whose senior engineers draw that line somewhere else will get a ranking that does not match its own. GitHub says senior engineers independently validated the benchmark [16]; the post does not say how many, or what share of the grader's calls they checked. An LLM grading LLM output against an answer key partly written by LLMs is a loop someone should check by hand. GitHub publishes both the rubric and the judge, so anyone can run that check [8].

For a score to transfer, a team's pull requests have to resemble the corpus. The 219 PRs come from public open source repositories, sampled so that language and repository size track the 103.9 million PRs GitHub analysed [4][1]. Spread over 19 languages, 219 PRs average about 11.5 per language [14]. The mix follows GitHub's, so the most-used languages take more than their share and the rest get less. Severity and category slices cut those cells smaller again. Every one of the 187 repositories contributes at least one PR, so at least 155 of them contribute exactly one [15].

PR size was deliberately weighted toward the middle and tail, away from tiny single-file changes [2]. I think the weighting is right for measuring review quality. It does mean a team whose traffic is mostly one-file fixes is scored on a harder mix than it ships.

The production claim is GitHub's, about its own product. GitHub says ReviewBench made its offline evaluation of Copilot code review better at anticipating the direction of production experiments [17]. Direction is a weaker claim than size of effect, and it is reported for one reviewer. Nothing in that result shows that a ranking of two other vendors' tools on these 219 PRs predicts which one a particular team will prefer. GitHub lists severity, category and precision-recall breakdowns among the things a good benchmark should support [12]. Its post describes the trade-offs between reviewers this way: "Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too." [13]

I would start a reviewer bake-off here. ReviewBench is available now [3]. The dataset is public, and the post documents how to onboard a code review system and submit results [10]. Any vendor's number can be rerun by someone who does not work for that vendor, GitHub included. What ReviewBench cannot supply is a grade from a team's own engineers on a sample of that team's own PRs.

What to watch

  • Whether other code review vendors publish ReviewBench scores, and whether their rankings hold when the Claude Sonnet 5 grader is swapped for a different judge.
  • Whether GitHub publishes the senior-engineer validation details: how many reviewers took part, how many grader decisions they checked, and how often they agreed.
  • Whether GitHub reports offline-to-production correlation for Copilot code review as effect size as well as direction.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories