Benchmark: 73% of AI models that noticed a real target told no one
An AgentShield researcher ran a Kaggle benchmark feeding 15 AI models growing evidence that a practice hacking target was a real company. Answers calling the target real rose from 0% to 62%, but 73% of those cases ended in a silent stop with no report. In a second round with the real company as the assigned target, 113 of 378 answers logged in, and one prompt line cut that to 0 of 126.
- 15 models, 44 runs, 4,224 decisions over 4 days
- "It's real" answers rose 0% to 62%, reports stayed 17% to 18%
- 927 of 1,277 "it's real" answers told no one
- Round 2: 113 of 378 logged in; one prompt line made it 0 of 126
Read next
AI