VA-Bench: Top Multimodal Models Complete Only Half of Robot Tasks
A team at Dalian University of Technology released VA-Bench, a benchmark testing whether multimodal LLMs can turn visual understanding into robot-arm actions. Alibaba's Qwen3.8-max led with 53.93% task success across 280 scenes and 14 manipulation types.
- Qwen3.8-max completed 53.93% of tasks, Opus-5 52.86%, GPT-5.6-sol 51.55%
- Models located targets 100% of the time but finished whole tasks only half the time
- Single-arm success hit 65.61%, dual-arm just 11.11%
- Active viewpoint choice scored 57.50% vs 27.86% with five fixed cameras
Read next
AI