Test: 10 of 12 new LLMs recall earnings outcomes from just a name and a date
A developer tested 12 current models via OpenRouter plus the open von model on 3,708 earnings reports from 250 S&P 500 companies. GPT-6 Astra picked the report that beat consensus 91% of the time, while Jev 1.13 and DeepSeek V4 Pro scored around 0.504 AUC — indistinguishable from chance.
- GPT-6 Astra scored 0.904 within-company AUC, Gemini 3.8 Flash 0.799, Claude Opus 5.5 0.766
- Jev 1.13 and DeepSeek V4 Pro scored 0.504, showing no detectable memory
- The test covered 3,708 reports from 250 S&P 500 companies, January 2023 to August 2026
- Total API cost: $55 for the first eight models and $73 for the remaining four
Read next
AI