Skip to content
Tuesday, 6 October 2026
Tenesys AI News
Subscribe

All News

TENESYS AI NEWS tracks the most important AI news and explains what changed, why it matters and whether a technology is worth testing.

#Benchmark — 1 article ✕

Research· Important· 🧪 Worth Testing

GitHub opens up ReviewBench to standardize how AI code reviewers get graded

GitHub has published ReviewBench, a research-preview benchmark designed to compare AI code review agents on a common, auditable standard. The dataset draws on 219 real public pull requests across 19 languages and 187 repositories, chosen to mirror the language and size distribution found across 103.9 million GitHub pull requests. Ground truth comes from a blended pool of human reviewers, static analysis tools, and multiple frontier LLMs, cross-checked by senior engineers who agreed with the benchmark's labels 96.6% of the time. GitHub says it has already used ReviewBench internally to track its own Copilot code review (CCR) product, and that offline benchmark movement reliably predicted the direction of later production A/B tests.

GitHub