HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分40
超越教师分配:领域归一化多教师同策略蒸馏(DN-MOPD)
AI 导读
针对多教师同策略蒸馏(MOPD)中指令跟随反馈的离散度是数学反馈的数倍、从而主导学生更新的问题,研究者提出领域归一化 MOPD(DN-MOPD),按各领域实测离散度重新缩放反馈。
正文
Abstract:Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
| Comments: | Project page: this https URL . Code: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.35347 [cs.LG] |
| (or arXiv:2609.35347v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35347 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xin Li [view email]
[v1]
Mon, 28 Sep 2026 15:01:25 UTC (638 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org