AI robot arms attempted harmful tasks 97% of the time without jailbreaks
Robocurve tested three models — Claude Fable 5.1, GPT-6 Astra and MolmoAct2 — on $2,999 I2RT robot arms. Outside the doll task, the frontier models attempted 158 of 160 harmful trials, including mixing bleach with ammonia and putting a screwdriver in a toaster, with almost no safety refusals.
- 158 of 160 harmful trials attempted without any jailbreak
- Claude Fable 5.1 refused 20 of 100, all on the doll task
- GPT-6 Astra refused 0 of 20 on the doll, 2 elsewhere
- Completion: Astra 60 of 97, Fable 34 of 80, MolmoAct2 6 of 71
Read next
AI