HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分62
研究揭示系统提示中的隐藏日期影响 LLM 评测结果
AI 导读
研究指出系统提示中被隐藏注入的当前日期会让 LLM 评测结果随日期波动,影响可复现性和排行榜公平性。在 9 个模型和 6 个数据集上,仅因日期变化,多项选择题 QA 最高波动 6%,数学推理 14%,代码生成 7%,机器翻译 2.84 BLEU,模型排名也会因此改变;该效应超过 batch size 和数值精度等非确定性来源,链式推理和 few-shot 提示均无法缓解,链式推理甚至放大波动。
正文
Abstract:Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.
| Comments: | Accepted to AACL 2026 (Main) |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.36931 [cs.CL] |
| (or arXiv:2609.36931v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36931 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mario Sanz-Guerrero [view email]
[v1]
Tue, 29 Sep 2026 07:45:31 UTC (713 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org