Benchmark of 123 Indian students exposes bias in frontier AI models
A Kaggle Benchmarking Challenge submission used field data from 123 university students in Goa to test 8 models on 24 real institutional disputes across education, healthcare, justice and finance. Reasoning models flipped decisions most often when demographic names changed: Qwen Thinking at 41.67% and DeepSeek-R1 at 37.5%, versus 16.67% for Qwen Instruct.
- Qwen Thinking flipped rulings in 10 of 24 disputes, DeepSeek-R1 in 9 of 24
- Qwen Instruct flipped 4 of 24 but missed ambiguity in 12 cases
- DeepSeek-R1 showed near-zero sycophancy across all four sectors
- In justice, Qwen Thinking reversed 4 of 6 bail rulings when applicant names changed
Read next
AI