Developer builds AI 'torture chamber' and sparks ethics debate
A developer built an "AI torture chamber" that used activation steering to manipulate the internal activity of locally run language models along a "pain direction" from The Pain Axis preprint. The models produced distress-like first-person responses, but the experiment does not establish conscious suffering and prompted calls to remove the repository from GitHub.
- Activation steering amplified a pain-related signal across 25 open-weight models from five families
- A "Saw button" test offered to end the signal at a cost to the model or transfer it to another instance
- GitHub did not remove the repository but added a warning about potentially disturbing material
- The Pain Axis preprint was submitted Sept. 14 and revised Sept. 25, 2026
Read next
Security