Creative Writing Benchmark: Frontier AI Matches Amateurs but Trails Professionals
A benchmark of 475 long-form prompts compared 24 LLMs with human writers: GPT 6 Astra scored an 87.8% predicted win rate versus 86.6% for the combined human group, though professional writers scored substantially higher when measured separately. Scoring relied on a reward model of human preference, not objective writing quality.
- GPT 6 Astra hit 87.8% predicted win rate, just above 86.6% for humans overall
- GPT 5.6 Sol scored 77.6%, Claude Fable 5.1 70.3%, while Claude Opus 5, Kimi K3 and Grok 4.6 landed at 50–65%
- Humans averaged 2592 tokens per prompt; Claude Fable 5.1 produced 2114 and GPT 6 Astra 1537
- Smaller models lose coherence and repeat themselves on longer prompts
Read next
AI