A code review benchmark that isn't the vendor ranking itself
AI lab Martian introduced Code Review Bench, an open benchmark for AI code review tools. On 16,017 real GitHub pull requests, Cubic Dev AI leads with F1 65.7%, followed by GitHub Copilot 63.9%, Claude 62.5%, CodeRabbit 60.8%, CodeAnt AI 50.0%. Methodology and code are published under the MIT license.
- Benchmark evaluated 14 tools on 16,017 GitHub pull requests
- Top F1: Cubic Dev AI at 65.7% with precision 72.9 and recall 59.8
- Greptile had best precision 80.3% but worst recall 50.4%
- Code and methodology are open on GitHub under MIT license
Read next
AI