HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分36
AlignOPSD:面向长程智能体的决策对齐在线策略蒸馏
AI 导读
针对长程智能体在线策略自蒸馏中的"决策—时间戳错配"问题,研究者提出 AlignOPSD,通过决策对齐监督校正与半马尔可夫分层信用分配,将监督对齐后再分配信用。
正文
Abstract:Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at this https URL
| Comments: | 27 pages, 10 figures |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.33391 [cs.LG] |
| (or arXiv:2609.33391v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33391 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mingju Chen [view email]
[v1]
Sun, 27 Sep 2026 09:13:27 UTC (1,410 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org