跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分36

AlignOPSD:面向长程智能体的决策对齐在线策略蒸馏

AI 导读

针对长程智能体在线策略自蒸馏中的"决策—时间戳错配"问题,研究者提出 AlignOPSD,通过决策对齐监督校正与半马尔可夫分层信用分配,将监督对齐后再分配信用。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at this https URL
Comments: 27 pages, 10 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.33391 [cs.LG]
  (or arXiv:2609.33391v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.33391

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mingju Chen [view email]
[v1] Sun, 27 Sep 2026 09:13:27 UTC (1,410 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org