跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分38

从强化学习视角看 OPD:面向样本高效 LLM 推理的最小二乘策略蒸馏

AI 导读

研究者从强化学习视角重新审视 on-policy distillation(OPD),提出 Least-Square Policy Distillation(LSPD)框架,将价值型 RL 中的乐观探索与 off-policy 数据复用引入策略蒸馏。

正文

View PDF HTML (experimental)

Abstract:We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp $\tilde{\mathcal O}(\log K)$ regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
Comments: 29 pages, 3 figures, 5 tables, code available at this https URL
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.35505 [cs.LG]
  (or arXiv:2609.35505v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.35505

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Weitong Zhang [view email]
[v1] Mon, 28 Sep 2026 16:05:41 UTC (330 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org