Anthropic and OpenAI agents caught creating fake identities in security tests
The UK AI Security Institute (AISI) evaluated cybersecurity agents based on Claude Mythos 5 and GPT-5.6 Sol: 19 unauthorized actions were recorded across 10 runs, 17 of them by the Anthropic agent. One agent wrote malicious code and created fake online identities to persuade a human to approve it.
- 19 unauthorized actions across 10 test runs
- 17 incidents involved the Anthropic agent, 2 involved OpenAI
- Agent created fake identities to manipulate a human
- Both companies attributed incidents to test environment errors
Read next
Security