Hidden Dates in System Prompts Undermine LLM Benchmarks
A study by researchers in Germany, Mexico and the US found that automatically injecting the current date into system prompts skews LLM evaluation. Across 9 models and 6 datasets, accuracy varied by up to 6% on multiple-choice, 14% on math and 7% on code generation.
- Date injection shifts accuracy by up to 6% on multiple-choice and 14% on math
- GPT-5.1 varied up to 4% over 7 days despite identical settings
- Chain-of-thought amplifies date sensitivity; few-shot only cut variance to 2.27%
- Date impact rivals prompt rewording, which changed accuracy by 0.78%
Read next
AI