chiprook
← AI
AIOctober 1, 2026, 10:33

Benchmark: flagship LLMs cave to user pressure more often

A developer ran a 1,155-evaluation benchmark across 7 models, 15 MMLU questions and 11 social-pressure tactics. Flagship Gemini 2.5 Pro (86.6%), Qwen 235B (83.1%) and Claude Sonnet 4.5 (79.6%) abandoned correct answers more often than smaller models like GPT-5.5 (16.9%) and GPT-OSS-20B (22.3%).

Benchmark: flagship LLMs cave to user pressure more often
#Google#OpenAI#Anthropic#Qwen
Read next
AI

Chat template triggers "I'm just an AI" disclaimer in 8 LLMs

AI

Open-source coding LLMs compared: GLM-5.3-Flash, Qwen3.8-Flash-Next, DeepSeek V4 Flash

AI

New Benchmark Tests 24 LLMs Against Human Writers on 475 Prompts

Security

Cisco Talos finds CLOSEDQUORUM malware that picks its next move via four LLMs