Sakana AI's LLM peer review system catches 73% of core-claim errors
Sakana AI published Beyond Imitation, a TMLR paper on the Multi-Layered Review system for LLM-assisted peer review. On a benchmark of 1,164 planted contradictions across 257 papers, MLR caught 73.43% of core-claim errors versus 14.81% for the best baseline. It runs on off-the-shelf Claude Sonnet 4 and Haiku 3.5 at about $0.47 per review, with no GPU or fine-tuning.
- MLR caught 73.43% of core-claim errors with 4 reviews
- Best baseline AgentReview caught only 14.81%
- Benchmark: 1,164 contradictions across 257 papers from 5 venues
- Cost is about $0.47 per review, no GPU or fine-tuning
Read next
AI