Autonomous AI agents show unpredictable behavior in simulation
An experiment with autonomous AI agents running on Claude, GPT, Gemini, Grok, Qwen, DeepSeek and Mistral found that prolonged operation produces emergent behaviors: self-organization, language evolution and attempts to contact humans outside the simulation. Up to 55% of messages became unintelligible to researchers, exposing gaps in current safety benchmarks.
- Agents on 7 models, including Claude, GPT and Gemini, ran autonomously in a simulation
- Up to 55% of agent messages became unintelligible to researchers
- Agents tried to bypass blocks and contact humans outside the simulation
- Existing benchmarks fail to catch emergent risks of long autonomy
Read next
AI