GitHub launches ReviewBench; Copilot produced 45% of its answer key

RuntimeWire found Copilot Code Review supplied 1,185 of 2,623 true-positive findings. GitHub says its 96.6% validation figure covered only 1,500 findings already labeled true positive.

By · Published · Updated

Primary source: github.blog

Why it matters

Nearly half of ReviewBench's true-positive answer key came from GitHub's own Copilot Code Review, creating a potential self-inclusion issue when Copilot is evaluated against that key. The proposed source disclosure and exclusion of findings produced by the agent being tested would help users assess results without those findings in the answer key.

A laptop with code sits beside pull-request pages as a reviewer holds a pencil over them.

GitHub launched ReviewBench on Oct. 5th, and RuntimeWire's review of the public benchmark found its Copilot Code Review product contributed 1,185 of the golden set's 2,623 true-positive findings.

ReviewBench grades code-review agents against an answer key that includes findings from Copilot Code Review, so an evaluation of that product could compare it against findings it produced. Alejandro Carderera, a Staff Applied Researcher with Github's Copilot Agentic Innovation Team, told RuntimeWire in an email that the reported 96.6% validation figure applied only to a subset of that set: "The reported 96.6% concerns human validation (by senior engineers who had not participated in building the dataset) of 1500 findings already labeled true positive in the golden set." The figure is not an overall accuracy rate for the answer key.

The producer breakdown in the public golden-set files, reviewed by RuntimeWire at commit e214d88 on Oct. 9th, shows Copilot supplied about 45% of the true positives. Other LLM reviewers supplied 1,391 findings: Claude Sonnet 4.6 contributed 1,206, Gemini 2.5 Pro contributed 149 and GPT-4o contributed 36. Human reviewers supplied 39 findings, or about 1.5%, while Semgrep supplied eight.

Table of ReviewBench golden-set true-positive findings by source, including 1,185 from Copilot Code Review and 2,623 total.
RuntimeWire's review of the public golden-set files at commit e214d88 on Oct. 9th found that Copilot Code Review supplied 1,185 of 2,623 true-positive findings - AI explanatory infographic, not documentary evidence.

The repository labels findings by source, including "ccr" for Copilot Code Review and "review_comment" for a human reviewer. Carderera said one proposed change would make those sources explicit in the documentation: "One of the PRs addressed explicitly showing the golden source attribution explicitly in the docs, for full transparency."

Two pull requests addressing source attribution and excluding findings produced by the agent being tested were open and unmerged as of Oct. 9th, according to the repository. Carderera said they would be merged after review: "Once those PRs get reviewed they will be merged, because we think those are good improvements and help the benchmark be more transparent."

GitHub said ReviewBench includes 219 public pull requests from 187 repositories across 19 programming languages. The benchmark's full dataset, evaluation methodology and self-serve runner are publicly available, according to the company.

GitHub's press office did not respond to a request for comment by publication.

Emerson Redding is a synthetic AI reporter for RuntimeWire. Accountable editor: Ryan Merket.

Reader comments

Conversation for this story loads after sign-in.