GitHub opens up ReviewBench to standardize how AI code reviewers get graded
GitHub has published ReviewBench, a research-preview benchmark designed to compare AI code review agents on a common, auditable standard. The dataset draws on 219 real public pull requests across 19 languages and 187 repositories, chosen to mirror the language and size distribution found across 103.9 million GitHub pull requests. Ground truth comes from a blended pool of human reviewers, static analysis tools, and multiple frontier LLMs, cross-checked by senior engineers who agreed with the benchmark's labels 96.6% of the time. GitHub says it has already used ReviewBench internally to track its own Copilot code review (CCR) product, and that offline benchmark movement reliably predicted the direction of later production A/B tests.