跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分62

研究揭示系统提示中的隐藏日期影响 LLM 评测结果

AI 导读

研究指出系统提示中被隐藏注入的当前日期会让 LLM 评测结果随日期波动,影响可复现性和排行榜公平性。在 9 个模型和 6 个数据集上,仅因日期变化,多项选择题 QA 最高波动 6%,数学推理 14%,代码生成 7%,机器翻译 2.84 BLEU,模型排名也会因此改变;该效应超过 batch size 和数值精度等非确定性来源,链式推理和 few-shot 提示均无法缓解,链式推理甚至放大波动。

正文

View PDF HTML (experimental)

Abstract:Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.
Comments: Accepted to AACL 2026 (Main)
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.36931 [cs.CL]
  (or arXiv:2609.36931v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.36931

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mario Sanz-Guerrero [view email]
[v1] Tue, 29 Sep 2026 07:45:31 UTC (713 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org