跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分32

PMOPD:多教师在线策略蒸馏中的任务排序、循环与参数更新子空间保护

AI 导读

针对多教师在线策略蒸馏(MOPD)中的能力跷跷板问题,研究者提出 PMOPD,通过从各任务累积参数位移构建子空间记忆,对梯度和优化器更新做投影以消除跨任务干扰,并配合轻量冲突探针指导任务排序与循环策略。在 Code、Reason、Math 任务上,PMOPD 全面优于 MOPD,Qwen2.5-7B 平均分提升 2.54 分,Llama-3.1-8B 提升 2.09 分。

正文

View PDF HTML (experimental)

Abstract:Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.34605 [cs.LG]
  (or arXiv:2609.34605v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.34605

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Youzhi Liu [view email]
[v1] Mon, 28 Sep 2026 08:36:16 UTC (2,664 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org