跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分47

同族同策略蒸馏的缩放规律:弱到强、同基座与强到弱教师-学生设置研究

AI 导读

研究揭示同策略蒸馏(OPD)在弱到强、同基座、强到弱教师-学生设置中的缩放规律:早期训练中,留出准确率(gold score)随学生初始化 token 级反向 KL 散度平方根近似线性上升,形成"有效迁移"区间。

正文

View PDF HTML (experimental)

Abstract:*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Comments: 35 pages, 20 figures
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.32722 [cs.LG]
  (or arXiv:2609.32722v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.32722

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yuntai Bao [view email]
[v1] Sat, 26 Sep 2026 15:39:11 UTC (2,130 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org