JetBrains ranks AI agents on real Kotlin projects, token use varies 12x
JetBrains introduced Kotlin Benchmark, an official benchmark for evaluating AI agents on real tasks in open Kotlin repositories. Claude Code with Opus 4.7 xhigh leads at 85.7%, but token spend per solved task varies up to 12 times, from 66,000 to 777,000.
- 105 tasks from active open-source Kotlin repositories, patches checked by repo tests
- Claude Code + Opus 4.7 xhigh: 90 tasks, 10.59 million tokens, 118,000 per task
- Claude Code + Opus 4.7 medium: 80 tasks at 66,000 tokens per task, half the leader's cost
- Gemini CLI + Gemini 3 Flash: 47 tasks and 777,000 tokens per task, worst result
Read next
AI