Microsoft and Hugging Face Launch ThinkingBox Benchmark for AI Agents
Microsoft and Hugging Face released ThinkingBox, a framework that grades AI agents by the actual database state they produce rather than their final reply. ThinkingBox-Bench spans 507 business workflows and 18 models: Claude Opus 5.5 leads pass@1 at 67.16%, while Kimi-K3 is the top open-weight model at 57.37%.
- ThinkingBox-Bench: 507 workflows across 5 domains, 20 runs per model, 18 LLMs
- 79,853 of 121,680 attempts failed checks though 67.24% ended with no tool error
- 77.61% of failures had wrong field values, 43.30% produced extra effects
- Claude Opus 5.5: 67.16% pass@1 but only 241 of 507 tasks passed all 20 runs
Read next
AI