Scale AI Reports ROK-FORTRESS Findings on Multilingual AI Safety
Scale AI and the Korean AI Safety Institute published results of the ROK-FORTRESS benchmark: Korean prompts produced lower measured harm in almost all 14 models. The gap between the safest and least safe model was nearly ninefold.
- Benchmark covers 1235 tasks in 4 domains: CBRNE, terrorism, crime, and data leaks
- Public set on Hugging Face: 791 tasks (64%), 444 (36%) held out privately
- Switching from English to Korean has 2.5 times stronger effect than changing geopolitical context
- In 5 open models, Korean prompts more often led to agreement to fulfill requests
Read next
AI