跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分42

用在线蒸馏缓解长度缩放税:Length Self-Distillation(LSD)方法

AI 导读

针对 RL 后训练中已解出问题回答变得冗长的"长度缩放税"(LST),研究者提出 Length Self-Distillation(LSD),将已解出的提示词路由到 on-policy 蒸馏,未解出的仍保留原 RL 目标,并以在线策略的指数移动平均作为教师,无需外部模型。

正文

View PDF HTML (experimental)

Abstract:Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants, while substantially curbing response-length growth on easy queries. LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks, demonstrating that LSD effectively preserves concise response patterns on easy queries while supporting efficient exploration on difficult queries during RL post-training.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.38854 [cs.LG]
  (or arXiv:2609.38854v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38854

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xu Wan [view email]
[v1] Wed, 30 Sep 2026 03:11:27 UTC (3,854 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org