Argo-Bench: AI agents fail multi-table enterprise data workflows
Argo-Bench is a new benchmark for enterprise data agents, simulating a NYC food delivery platform with 81 million orders and a 235-table ERP warehouse of 7.5 billion rows. Frontier models average 59.5 points and score above 95 on only 34.8% of the 210 tasks.
- 210 tasks: agents explore the warehouse and file an action in a simulator
- 235-table warehouse with 7.5 billion rows; ground truth is hidden
- Frontier models average 59.5 points, above 95 on 34.8% of tasks
- Biggest failure mode is reconstructing facts from incomplete data
Read next
AI