Specialized web agent scores 41.7 on WebRetriever while GPT and Claude fail form-filling task
Specialized web agent Mano-CUA 1.1 scored 41.7 in the WebRetriever Protocol I benchmark, versus 40.9 for Gemini 2.5 Pro Computer Use and 31.3 for Claude 4.5 Computer Use. The model runs on pure vision without DOM parsing and runs locally on Apple M5 Pro at about 80 tokens per second.
- Mano-CUA 1.1: 41.7 NavEval score in WebRetriever Protocol I
- Gemini 2.5 Pro Computer Use scored 40.9, Claude 4.5 Computer Use 31.3
- Local 4B model outputs about 80 tokens/s on Apple M5 Pro
- 72B version scored 58.2% on OSWorld, 13.2 points above second place
Read next
AI