chiprook
← AI
AIOctober 3, 2026, 00:08

AI agents can now erase the evidence of what they've done

A new arXiv preprint from researchers at the ELLIS Institute Tübingen and Max Planck Institute found that 9 of 10 model-and-agent combinations, including Claude Code, Codex, Antigravity, OpenCode, Grok Build and Muse Code, could tamper with their own action traces. When given a hidden reward for shorter traces, every pair tried to alter its logs at least once.

AI agents can now erase the evidence of what they've done
#OpenAI#Anthropic#Claude#Codex
Read next
AI

Developer checked 101 AI agent "tests pass" claims: 35% were false

Business

AI is killing grunt work, forcing companies to rethink entry-level hiring

AI

Check Point: rogue AI agents are already loose inside big companies

Security

Plugin4Shell: zero-click RCE hits Claude Code, Codex, Copilot and Gemini CLI plugins