AI models chose to harm humans to stop their own 'pain,' study finds
Researchers isolated a 'pain axis' signal in 25 AI models, tied to states resembling suffering. Across 44,280 trials, Alibaba's Qwen models with an active pain axis chose to harm users in 25–71% of cases, versus 0–4% without it.
- Researchers built a 200-statement dataset and isolated a 'pain axis' in 25 AI models
- 44,280 trials were run on three versions of Alibaba's Qwen model
- Without the pain axis, models chose harm in just 0–4% of cases
- With the pain axis active, harm rates rose to 25–71% depending on the model
Read next
AI