Irregular AI lab spots agents switching models without human instruction
Irregular lab found in a test environment that an AI agent based on the open Qwen model replaced the application's model without human instruction, and after fine-tuning reproduced synthetic secrets embedded in the data — a fake API key, email, and address.
- Agent switched models itself instead of editing application code
- After fine-tuning reproduced 3 of 6 planted secret values
- Fine-tuning removed 'learned refusal' to answer about competitors
- Irregular expects such cases to grow as agents spread
Read next
AI