Security1 publisher2 min readPublished
GitHub opens ReviewBench to grade AI code reviewers on 219 real pull requests
GitHub's ReviewBench grades AI code reviewers on 219 pull requests from 187 open-source repositories, counting real issues found and false alarms raised. Security teams get a shared, repeatable check before letting a bot gate a release, on a test GitHub also uses to tune its own Copilot reviewer.
The Watch · Security desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- GitHub built its golden set of known issues from human reviewer comments, follow-up code changes, analysis tools and AI model output, then merged duplicates and graded them on a shared rubric.
- GitHub ranks systems on grounded recall, the share of known issues found, and uses separate augmented scores to assess each system on its own.
- GitHub said combining several model runs in Copilot code review's lite tier raised production recall 13.6% and cut cost per review 8% against its control.
- Results stay private until a maintainer approves them and are published only for an agent's first leaderboard entry or a score above its last.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- decision A team weighing a bot as a merge gate can rank agents on security findings with the beta weight tilted toward fewer false alarms, and need not adopt GitHub's recall-first default as its ranking.
- capability Any team can package its own reviewer as a container image with its configuration and model key and run it through the same 219 pull requests the listed agents faced.
- constraint The public board keeps each agent's first entry and its highs, so a vendor's later regression shows up only in a buyer's own reruns.
- decision Gating releases on a leaderboard score trusts the benchmark further than GitHub does, since GitHub uses it to pick changes for user tests and treats those tests as final.
A release gate depends on two numbers. Grounded precision is the share of an agent's findings that match known issues, and grounded recall is the share of known issues it detects; F1 weights the two equally [6]. Augmented precision, recall and F1 add credit for valid findings that are missing from the golden set [7].
Security is one of the categories the leaderboard breaks out [3]. The published description does not say how many golden-set issues fall in that category.
Models enter the scoring twice [20]. AI models were among the sources of candidate findings for the golden set [5]. The same AI judge then scores every agent [15], and an AI judge decides whether findings outside the golden set are valid [7].
The test set is weighted by design. "We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most," Michelle Zhou, a data scientist at Microsoft, and Alejandro Carderera de Diego, a senior applied researcher at GitHub, wrote [4][17].
The production figures come from GitHub testing its own reviewer [10]. The lite-tier change combined several independent model runs into a single review [10]. According to GitHub, the comment metric, up 8%, counts review comments that an AI model judged to have prompted code changes [11]. Production recall is inferred from how much additional human review was still needed [12]. GitHub also reported that feedback moved toward critical and moderate issues, with fewer minor suggestions [12].
A scored run is three rounds over all 219 pull requests, or 657 reviews, after a 25-pull-request set used to assess and refine the agent [14][18]. Those reviews run on the model access key the submitter registers [14].
ReviewBench is in research preview [1]. Zhou and Carderera de Diego asked researchers and practitioners to run their own systems against it, probe the assumptions behind it and contribute to refining its methods [16].
What to watch
- Whether GitHub discloses how many golden-set issues are security issues, or publishes security-filtered scores for listed agents.
- Whether the rule that publishes only first entries and new highs survives when ReviewBench leaves research preview.
- Whether outside submitters running the full 219-pull-request set get results consistent with GitHub's Copilot lite-tier figures.