Safe Labs: an AI agent's refusal doesn't prove its actions were safe
Safe Labs AI released safelabs-trace, a benchmark comparing text-only and action-aware safety evaluation across 9,900 agent runs, 300 adversarial tasks, six models and three frameworks. Risky tool calls appeared in 7.6% of low-cost and 3.0% of frontier trials, while the text scorer abstained on 43.2% and 31.1% of runs.
- 9,900 runs: 300 tasks, 6 models, 3 frameworks, 12 tools
- Risky actions flagged in 7.6% of low-cost and 3.0% of frontier trials
- Tool-call evidence raised detection from 3.5% to 4.7% and 2.7% to 3.7%
- Text scorer abstained on 43.2% and 31.1% of trials
Read next
AI