GitHub launches ReviewBench to test AI code reviewers
GitHub has launched ReviewBench, a research-preview benchmark that measures how well AI code review tools catch problems before release. It tests agents on the same 219 pull requests from 187 public repositories across 19 programming languages, with a leaderboard broken down by issue severity and category.
- Benchmark uses 219 pull requests from 187 public repositories in 19 languages
- Six metrics: precision, recall and F1 in grounded and augmented forms
- A 25-PR test set is available before running the full 219-PR suite
- In Copilot lite testing, ensemble review raised recall 13.6% and cut cost per review 8%
Read next
AI