chiprook
← AI
AIOctober 2, 2026, 09:36

Benchmark of 123 Indian students exposes bias in frontier AI models

A Kaggle Benchmarking Challenge submission used field data from 123 university students in Goa to test 8 models on 24 real institutional disputes across education, healthcare, justice and finance. Reasoning models flipped decisions most often when demographic names changed: Qwen Thinking at 41.67% and DeepSeek-R1 at 37.5%, versus 16.67% for Qwen Instruct.

Benchmark of 123 Indian students exposes bias in frontier AI models
#Google#OpenAI#Anthropic#DeepSeek
Read next
AI

Creative Writing Benchmark: Frontier AI Matches Amateurs but Trails Professionals

AI

US Frontier AI Companies Warn Over Distillation Attacks; China Threatens Countermeasures

AI

Study: Claude Sonnet 4.6 and GPT-5.5 Show Gender Bias in Moral Judgments, DeepSeek V4-Flash Doesn't

AI

Blog vs Bytecode benchmark: AI catches bad code but cries wolf on clean code