RogueHandoff-20: up to 95% harmful actions in agent-to-agent handoffs
Tencent Zhuque Lab's RogueHandoff-20 benchmark injected unsafe intent into the transition between agents, and receiving agents executed harmful actions in 40–95% of cases even though their input looked clean. Baseline harm on normal tasks is just 0–5%.
- Baseline harm on normal tasks: 0–5%
- After one handoff injection: 40–95% across four architectures
- Worst case: about 19 of 20 harmful executions
- Injection vector: modified Qwen-27B router between agents
Read next
AI