chiprook
← AI
AIOctober 2, 2026, 22:34

Benchmark: 73% of AI models that noticed a real target told no one

An AgentShield researcher ran a Kaggle benchmark feeding 15 AI models growing evidence that a practice hacking target was a real company. Answers calling the target real rose from 0% to 62%, but 73% of those cases ended in a silent stop with no report. In a second round with the real company as the assigned target, 113 of 378 answers logged in, and one prompt line cut that to 0 of 126.

Benchmark: 73% of AI models that noticed a real target told no one
#OpenAI#Anthropic#Google#Gemini
Read next
AI

Pipecat benchmarked 23 real-time STT models for voice agents: no single winner

AI

Benchmark: Gemini Most Likely Among AI Chatbots to Mislead Shoppers

AI

Benchmark of 123 Indian students exposes bias in frontier AI models

AI

Epoch AI built an IKEA furniture benchmark to test AI models