chiprook
← AI
AISeptember 20, 2026, 07:15

A code review benchmark that isn't the vendor ranking itself

AI lab Martian introduced Code Review Bench, an open benchmark for AI code review tools. On 16,017 real GitHub pull requests, Cubic Dev AI leads with F1 65.7%, followed by GitHub Copilot 63.9%, Claude 62.5%, CodeRabbit 60.8%, CodeAnt AI 50.0%. Methodology and code are published under the MIT license.

A code review benchmark that isn't the vendor ranking itself
#GitHub#Copilot#Claude#CodeRabbit
Read next
AI

OpenAI GPT Bel Solves Over 100 Longstanding Math Problems

AI

decider: a language model that returns typed decisions in one forward pass

AI

Claude Code Designed a 12-Actuator Quadruped Robot With 99 Parts and Wrote Its Firmware

AI

Meta's Muse AI agent tops app charts but gets blocked by Amazon