跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分36

DCSD:解耦信用方向与幅度的自蒸馏方法,11 项基准超越 GRPO 与 OPSD

AI 导读

针对教师监督下步级信用方向与幅度耦合的问题,研究者提出解耦信用自蒸馏(DCSD),用 belief-margin probing 确定信用方向、用边际信息增益量化信用幅度,实现步到 token 的信用分配。

正文

View PDF HTML (experimental)

Abstract:RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6\% of tokens and yielding a 1.5$\times$ reduction in token credit magnitude.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.34848 [cs.AI]
  (or arXiv:2609.34848v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.34848

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yugu Li [view email]
[v1] Mon, 28 Sep 2026 10:43:13 UTC (4,363 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org