AI chatbots get 57% of financial questions wrong, study finds
Saturn tested 18 popular AI tools, including ChatGPT, Gemini, Claude and Copilot: average accuracy on financial questions was 43%, dropping to 12% on complex multi-step scenarios. Paid Claude Opus 5 (reasoning) performed best with a 61% pass rate, while free Claude Haiku 4.5 was worst with 82% wrong answers.
- Average accuracy across 17 models was 43%, wrong 57% of the time
- Accuracy fell to 12% on complex multi-step scenarios
- Free models failed 63% of the time, paid models 49%
- Best performer was paid Claude Opus 5 (reasoning) at 61%
Read next
AI