KaliBench benchmark exposes AI agents' inability to build correct Kali Linux tool commands
KaliBench, a new benchmark of 8,504 query-command pairs covering 1,642 Kali Linux tools across 23 capability dimensions and 5 security phases, measures command generation without executing dangerous tools. No open model exceeded 42% exact-command accuracy without hints.
- 8,504 query-command pairs across 1,642 Kali Linux tools
- Best open model hit only 42% exact-command accuracy without hints
- Accuracy rises to 68% with tool hints and 71% with few-shot examples
- Verification uses canonicalization, LLM validation, sandboxing and human review
Read next
AI