chiprook
← AI
AIOctober 9, 2026, 21:37

Hidden Dates in System Prompts Undermine LLM Benchmarks

A study by researchers in Germany, Mexico and the US found that automatically injecting the current date into system prompts skews LLM evaluation. Across 9 models and 6 datasets, accuracy varied by up to 6% on multiple-choice, 14% on math and 7% on code generation.

Hidden Dates in System Prompts Undermine LLM Benchmarks
#OpenAI#GPT-5.1#Llama
Read next
AI

REAL-Q uses dynamic gradient descent to fix LLM quantization

Software

iOS 27.2 Beta: Hidden Features and October 19 Release Date

Software

llama-server prompt cache reuses KV computed under a different LoRA scale

AI

New Benchmark Tests 24 LLMs Against Human Writers on 475 Prompts