A joint study by researchers from Germany, Mexico and the USA found that automatic injection of the current date into system prompts undermines LLM evaluation. Across 9 models, 6 datasets and four tasks, this hidden metadata changed performance and reshuffled leaderboard rankings. The effect exceeded variance from batch size or numerical precision, which matters for buyers relying on benchmarks to compare vendors.

Hidden Dates in System Prompts Distort LLM Benchmark Results

How the hidden date changes test scores

The team tested every day of 2024, from January 1 to December 31, keeping user prompts and settings identical and changing only the date in the system prompt. Time-dependent questions were excluded. Six benchmarks were used: MMLU, GPQA and ARC-Challenge for multiple choice, GSM8K for math, HumanEval for Python code generation, and WMT for English-to-German, Finnish and Czech translation. Nine open-weight models were covered, including Llama 3.1 Instruct 8B and 70B, Gemma 3 Instruct 4B and 27B, Qwen3 4B, Qwen3-Next 80B, Phi-4 14B and GPT-OSS 20B and 120B.

Simply changing the date altered accuracy on identical questions. Variation reached up to 6% on multiple-choice questions, 14% on mathematical reasoning and 7% on code generation, with machine translation scores also shifting significantly. Accuracy measured correct answers, Expected Calibration Error measured confidence alignment, code used pass@1 and translation used BLEU and chrF. Generated-answer tasks proved more sensitive than multiple choice scored by answer-token probability. A model could therefore rise or fall on a leaderboard without any real change in capability.

The pattern held for proprietary systems where users cannot control the prompt. GPT-5.1 was tested on three multiple-choice benchmarks over seven consecutive days in December 2025 with an empty system prompt, disabled reasoning and randomness set to zero. Accuracy still fluctuated by up to 4%, with the largest gap on GPQA. The provider inserted the date automatically despite the fixed configuration. Chain-of-thought prompting amplified the problem, and few-shot learning with five examples cut average variation only from 2.52% to 2.27%.

What unstable benchmarks mean for business

For companies selecting models, date-driven noise weakens the signal from public leaderboards and vendor scorecards. A gap of several percentage points may reflect the test day rather than genuine progress between versions or competitors. Procurement teams comparing sales assistants, support agents or coding copilots risk overpaying for a lead that disappears on rerun. Small firms that rely on published rankings face the highest exposure, while large firms running internal pilots can average results across days but spend more compute and staff time.

The study points to prompt brittleness as a likely cause. The authors suggest the system prompt may be fixed during supervised fine-tuning and reinforcement learning from human feedback, making models sensitive to slight modifications. Six alternative system-prompt wordings affected accuracy by 0.78%, about as much as changing the date. Hardware, batch size, precision, answer order and wording were also tested, with answer order and wording closest to the date effect. Their recommendation is to remove the date where possible or fix and document it.

The practical marker is documentation of evaluation conditions. Businesses should ask vendors for the exact system prompt, the fixed date used during testing, and scores averaged over multiple dates rather than a single run. Contracts for agents and automation can require reruns on the buyer dataset with disclosed settings. If providers start publishing fixed-date, reproducible benchmarks and date-averaged results, selection risk will fall; until then, leaderboard differences of a few points should not drive purchasing decisions.