Skip to content

Security1 publisher2 min readPublished

GitHub opens ReviewBench to grade AI code reviewers on 219 real pull requests

GitHub's ReviewBench grades AI code reviewers on 219 pull requests from 187 open-source repositories, counting real issues found and false alarms raised. Security teams get a shared, repeatable check before letting a bot gate a release, on a test GitHub also uses to tune its own Copilot reviewer.

The Watch · Security desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying GitHub opens ReviewBench to grade AI code reviewers on 219 real pull requests
Generated illustration

What happened

  • GitHub built its golden set of known issues from human reviewer comments, follow-up code changes, analysis tools and AI model output, then merged duplicates and graded them on a shared rubric.
  • GitHub ranks systems on grounded recall, the share of known issues found, and uses separate augmented scores to assess each system on its own.
  • GitHub said combining several model runs in Copilot code review's lite tier raised production recall 13.6% and cut cost per review 8% against its control.
  • Results stay private until a maintainer approves them and are published only for an agent's first leaderboard entry or a score above its last.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • decision A team weighing a bot as a merge gate can rank agents on security findings with the beta weight tilted toward fewer false alarms, and need not adopt GitHub's recall-first default as its ranking.
  • capability Any team can package its own reviewer as a container image with its configuration and model key and run it through the same 219 pull requests the listed agents faced.
  • constraint The public board keeps each agent's first entry and its highs, so a vendor's later regression shows up only in a buyer's own reruns.
  • decision Gating releases on a leaderboard score trusts the benchmark further than GitHub does, since GitHub uses it to pick changes for user tests and treats those tests as final.

A release gate depends on two numbers. Grounded precision is the share of an agent's findings that match known issues, and grounded recall is the share of known issues it detects; F1 weights the two equally [6]. Augmented precision, recall and F1 add credit for valid findings that are missing from the golden set [7].

Security is one of the categories the leaderboard breaks out [3]. The published description does not say how many golden-set issues fall in that category.

Models enter the scoring twice [20]. AI models were among the sources of candidate findings for the golden set [5]. The same AI judge then scores every agent [15], and an AI judge decides whether findings outside the golden set are valid [7].

The test set is weighted by design. "We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most," Michelle Zhou, a data scientist at Microsoft, and Alejandro Carderera de Diego, a senior applied researcher at GitHub, wrote [4][17].

The production figures come from GitHub testing its own reviewer [10]. The lite-tier change combined several independent model runs into a single review [10]. According to GitHub, the comment metric, up 8%, counts review comments that an AI model judged to have prompted code changes [11]. Production recall is inferred from how much additional human review was still needed [12]. GitHub also reported that feedback moved toward critical and moderate issues, with fewer minor suggestions [12].

A scored run is three rounds over all 219 pull requests, or 657 reviews, after a 25-pull-request set used to assess and refine the agent [14][18]. Those reviews run on the model access key the submitter registers [14].

ReviewBench is in research preview [1]. Zhou and Carderera de Diego asked researchers and practitioners to run their own systems against it, probe the assumptions behind it and contribute to refining its methods [16].

What to watch

  • Whether GitHub discloses how many golden-set issues are security issues, or publishes security-filtered scores for listed agents.
  • Whether the rule that publishes only first entries and new highs survives when ReviewBench leaves research preview.
  • Whether outside submitters running the full 219-pull-request set get results consistent with GitHub's Copilot lite-tier figures.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories