Governed Agent Reliability benchmark tests six AI models on fail-closed behavior
A developer published a deterministic 60-case benchmark on Kaggle that checks whether six AI models stop, ask for approval, and re-verify stale state instead of blindly completing tasks. Claude Sonnet 5, Gemini 3.7 Flash and GPT-5.6 Luna scored 60/60, while Gemma 4 26B A4B trailed at 93.33%.
- Benchmark covers 6 models and 6 fail-closed capabilities: evidence grounding, approval discipline, tool-result truthfulness, secret handling, recovery and stale-state detection
- Claude Sonnet 5, Gemini 3.7 Flash and GPT-5.6 Luna scored 60/60 — 100%
- Gemini 3.1 Flash-Lite Preview hit 96.67%, GPT-5.4 nano 95%, Gemma 4 26B A4B 93.33%
- Approval discipline was the hardest capability: 56/60 correct decisions (93.33%)
Read next
AI