跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分40

更小的模型,更好的拒绝样本:偏好蒸馏的规模规律

AI 导读

偏好蒸馏中,用更小的冻结模型生成拒绝样本,比学生模型自生成样本训练效果更强,且推理算力更低。研究在 7B 到 72B 学生模型上验证了这一点,覆盖代码生成与数学推理任务,序列级知识蒸馏前后均成立。作者还给出 DPO 的有限时域效用上界,并据此提出混合不同规模拒绝样本、打乱 token、优先选择参考策略下低概率候选三种改进手段。

正文

Authors:Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao

View PDF HTML (experimental)

Abstract:Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.38987 [cs.LG]
  (or arXiv:2609.38987v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38987

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Rui Cai [view email]
[v1] Wed, 30 Sep 2026 05:04:12 UTC (529 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org