跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分34

DuoOPD:用师生联合结果做多任务在线策略蒸馏

AI 导读

DuoOPD 提出一种多任务在线策略蒸馏方法,由学生结果决定反馈方向、师生联合结果决定教师如何支持:仅教师答对时用其答案作为评分上下文,仅学生答对时按任务内共享权重强化整段回答。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.33711 [cs.LG]
  (or arXiv:2609.33711v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.33711

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ao Yu [view email]
[v1] Sun, 27 Sep 2026 16:15:02 UTC (852 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org