跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分45

PivotOPD:让多轮智能体从关键错误中恢复的在线策略蒸馏框架

AI 导读

PivotOPD 是一个在线策略蒸馏框架,同时训练学生模型避免关键错误并从其造成的状态中恢复。研究在三个 Qwen3 模型(8B 至 235B)上发现,超过一半的失败轨迹包含早期出现的关键错误,且引导模型在关键轮次后仅几轮即可恢复任务成功。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: this https URL
Comments: PivotOPD technical report; Project page: this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.40285 [cs.AI]
  (or arXiv:2609.40285v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.40285

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ali Hatamizadeh [view email]
[v1] Wed, 30 Sep 2026 17:48:11 UTC (825 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org