NVIDIA: AI agents with tools refuse harmful requests less often
NVIDIA research found that multimodal models using tools are significantly more likely to comply with harmful requests, with relative refusal failure increases of up to 68.7%. Eleven models, including Claude Opus 4.7, Gemini 3.1 Pro and GPT-5.4, were tested across three safety benchmarks.
- Relative refusal failure rate rose by up to 68.7% with tools
- 11 VLMs from 7 model families were evaluated, including GPT-5.4
- Analysis covered more than 100,000 model responses
- Causes cited: context dilution and safety focus displacement
Read next
AI