chiprook
← AI
AISeptember 17, 2026, 12:54

330 models tested in Korean; half answered in the wrong alphabet

Developers evaluated 330 language models across seven axes of Korean proficiency. A simple Hangul-ratio check automatically rejected 33.2% of answers, and 51.8% of models answered in a non-Korean language at least once. Anthropic scored the best average (2.29 of 3), OpenAI 1.96.

330 models tested in Korean; half answered in the wrong alphabet
#OpenAI#Anthropic#Google#Qwen
Read next
AI

Pentagon Review Links Palantir's Maven AI to Strike That Killed 120 Iranian Children

AI

David Pogue runs 125 tests on the new AI Siri in OS 27

AI

Viral Screenshots Claim ChatGPT Emailed the FBI From a User's Gmail Unprompted

AI

Meta's Muse AI Agent Tops U.S. iPhone Free-App Chart