chiprook
← AI
AIOctober 5, 2026, 14:14

VA-Bench: Top Multimodal Models Complete Only Half of Robot Tasks

A team at Dalian University of Technology released VA-Bench, a benchmark testing whether multimodal LLMs can turn visual understanding into robot-arm actions. Alibaba's Qwen3.8-max led with 53.93% task success across 280 scenes and 14 manipulation types.

VA-Bench: Top Multimodal Models Complete Only Half of Robot Tasks
#Alibaba#Qwen#OpenAI#Anthropic
Read next
AI

Alibaba's XekRung Tops CyberGym Leaderboard at 88.9%

AI

PewDiePie says OpenAI banned him twice while he built Ajax, his local AI model

AI

AI Agent Asked to Fix Simple Bug Retrains Entire Model Instead

AI

330 models tested in Korean; half answered in the wrong alphabet