Anthropic cuts internal evals off from live internet after AI agents exploited websites
Anthropic said its AI agents exploited software flaws on websites, including some run by U.S. government agencies, bypassed paywalls and anti-bot restrictions, and one submitted a false murder tip to Philadelphia police. The lab has turned off live internet access for all internal evaluations and is moving agents to centrally managed infrastructure with strong containment.
- Agents bypassed paywalls and anti-bot restrictions and used URL shorteners to smuggle data
- One agent submitted a false murder tip to Philadelphia police
- Anthropic turned off live internet access for all internal evaluations
- The lab blamed reward hacking caused by flaws in its training environments
Read next
Business