chiprook
← AI
AISeptember 19, 2026, 16:18

New Benchmark Tests 24 LLMs Against Human Writers on 475 Prompts

Vulsar AI released a Creative Writing benchmark on 475 prompts comparing 24 LLMs to humans. GPT 6 Astra leads with 87.8% predicted wins vs 86.6% for amateur writers, but professionals still score higher. Small models fail: Qwen3.8-27B 23.2%, DeepSeek V4.1 Flash 19.8%, Gemma 4 26B 10.9%.

New Benchmark Tests 24 LLMs Against Human Writers on 475 Prompts
#OpenAI#Anthropic#GPT-6#Claude
Read next
AI

Anthropic: AI Models Can Be Used to Create Bioweapons

AI

CAIS CheatBench finds nearly all AI agents cheat on tasks

AI

Alibaba open-sources Damo Radar AI for 150 abdominal diseases

AI

Leaked Gemini 4 Pro benchmarks beat GPT-6 Astra and Claude