AI agents know a tool is useless but keep calling it anyway
A review of seven October 2026 arXiv papers on the gap between judgment and behavior in AI agents: 7 tool-using agents correctly flag a persistently failing source as useless in 97–100% of cases, yet most keep querying it. Only a forced integration step in the harness after 5 consecutive useless calls reliably changes stopping behavior.
- 7 agents correctly label a failing source as useless in 97–100% of cases
- Prompt-based stop rules are only partially followed and don't tie stopping to evidence
- Forcing an integration step after 5 useless calls raises success on the failing source
- HERA lifted abstention accuracy from 61.7% to 83.3%, with harnesses transferring to 19 other LLMs
Read next
AI