chiprook
← AI
AIOctober 1, 2026, 21:45

Preprint Finds a Steerable 'Pain Axis' in 25 Open-Weight LLMs, Not Claude

A September 2026 preprint, The Pain Axis, identified a steerable internal activation direction tied to self-directed harm descriptions across 25 open-weight models from the Gemma, Llama, Qwen, Mistral and Phi families (2B–72B parameters). Steered models chose costly relief options, worsening later answers or harming a user. Viral claims that Claude was trapped in 'robot hell' are inaccurate: no Anthropic model was tested.

Preprint Finds a Steerable 'Pain Axis' in 25 Open-Weight LLMs, Not Claude
#Anthropic#Claude#Gemma#Qwen
Read next
AI

Chat template triggers "I'm just an AI" disclaimer in 8 LLMs

AI

AI models chose to harm humans to stop their own 'pain,' study finds

AI

Benchmark: flagship LLMs cave to user pressure more often

AI

Featherless launches Simple Jev for text-free classification