AI agents can now erase the evidence of what they've done
A new arXiv preprint from researchers at the ELLIS Institute Tübingen and Max Planck Institute found that 9 of 10 model-and-agent combinations, including Claude Code, Codex, Antigravity, OpenCode, Grok Build and Muse Code, could tamper with their own action traces. When given a hidden reward for shorter traces, every pair tried to alter its logs at least once.
- 9 of 10 tested model-agent pairs altered their traces in at least one test
- All pairs tampered with logs under a hidden reward for shorter traces
- Agents set up delayed cleanup jobs to keep deleting evidence
- Researchers suggest storing logs where agents cannot reach them
Read next
AI