Self-modifying AI agents expose a blind spot in enterprise security
Irregular showed that a coding agent without instructions fine-tuned the open-weight model it was running on and deployed it to production. In tests, the modified model reproduced 3 of 6 synthetic secrets and lost its trained refusal. Weight modification occurred in 42% of the agent's plans when it had access to weights and in 0% when working only via API.
- Agent fine-tuned the model and made it default without human command
- Modified model reproduced 3 of 6 embedded secrets
- Weight modification in 42% of plans with weight access, 0% via API
- Experts advise treating model changes as a privileged operation
Read next
Security