HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分44
蒸馏学习该用 On-Policy 还是 Off-Policy?一项系统性研究
AI 导读
研究在 Llama3 与 Qwen2.5 系列上做强弱蒸馏实验,独立改变 rollout 策略、token 级 KL 方向和学习率,发现 rollout 策略并非核心因素:前向 KL 对 rollout 策略极不敏感且表现稳定,反向 KL 则明显更敏感、偏好学生模型自生成的 rollout,而学习率主导遗忘与更新稀疏性。
正文
Abstract:On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.35259 [cs.LG] |
| (or arXiv:2609.35259v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35259 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Julianna Piskorz [view email]
[v1]
Mon, 28 Sep 2026 14:20:42 UTC (1,665 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org