chiprook
← AI
AISeptember 25, 2026, 20:00

Test: 10 of 12 new LLMs recall earnings outcomes from just a name and a date

A developer tested 12 current models via OpenRouter plus the open von model on 3,708 earnings reports from 250 S&P 500 companies. GPT-6 Astra picked the report that beat consensus 91% of the time, while Jev 1.13 and DeepSeek V4 Pro scored around 0.504 AUC — indistinguishable from chance.

Test: 10 of 12 new LLMs recall earnings outcomes from just a name and a date
#OpenAI#Anthropic#Google#DeepSeek
Read next
AI

Open-source coding LLMs compared: GLM-5.3-Flash, Qwen3.8-Flash-Next, DeepSeek V4 Flash

Security

Cisco Talos finds CLOSEDQUORUM malware that picks its next move via four LLMs

AI

New Benchmark Tests 24 LLMs Against Human Writers on 475 Prompts

AI

An AI meant to learn from its mistakes exploited a mistake in the test