330 models tested in Korean; half answered in the wrong alphabet
Developers evaluated 330 language models across seven axes of Korean proficiency. A simple Hangul-ratio check automatically rejected 33.2% of answers, and 51.8% of models answered in a non-Korean language at least once. Anthropic scored the best average (2.29 of 3), OpenAI 1.96.
- Auto-check rejected 766 of 2304 answers (33.2%) for Chinese or Japanese characters
- 171 of 330 models (51.8%) failed the check at least once, 54 failed all seven tasks
- Worst areas: honorifics (8.5% A grades) and Korean institutions (9.4%)
- Average scores: Anthropic 2.29, Mistral 2.07, OpenAI 1.96, Google and Qwen 1.48 each
Read next
AI