chiprook
← AI
AIOctober 11, 2026, 02:46

Safe Labs: an AI agent's refusal doesn't prove its actions were safe

Safe Labs AI released safelabs-trace, a benchmark comparing text-only and action-aware safety evaluation across 9,900 agent runs, 300 adversarial tasks, six models and three frameworks. Risky tool calls appeared in 7.6% of low-cost and 3.0% of frontier trials, while the text scorer abstained on 43.2% and 31.1% of runs.

Safe Labs: an AI agent's refusal doesn't prove its actions were safe
#SafeLabs
Read next
AI

Safeworld raises $12M to prove gen AI robots are safe

AI

Okta turns Dex AI agent into a customer-zero proving ground

AI

NVIDIA: AI agents with tools refuse harmful requests less often

AI

Robinhood CEO Says New AI-Driven Agent Apps Are Safe