跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-05-04精选AI 评分70

T^2PO:面向稳定多轮智能体强化学习的不确定性引导探索控制框架

AI 导读

多轮强化学习训练常因探索效率低下而不稳定。为此,研究团队提出T^2PO框架,在细粒度层面实施不确定性引导的探索控制。在令牌级别,它监测不确定性动态,当边际变化低于阈值时触发思考干预;在轮次级别,它识别探索进展可忽略的交互并动态重采样,以避免无效计算。在WebShop、ALFWorld和Search QA等多个环境中的评估表明,T^2PO显著提升了训练稳定性与任务性能,并实现了更高效的探索。相关代码已开源。

推荐理由

多轮 agent 训练最怕训着训着崩了,这篇从 token 和 turn 两级控制探索的思路很妙,直接把低效 rollout 砍掉,稳定性和效率都上去了,做 RLHF 或 agent RL 的可以认真看一下。

正文

View PDF HTML (experimental)

Abstract:Recent progress in multi-turn reinforcement learning (RL) has significantly improved reasoning LLMs' performances on complex interactive tasks. Despite advances in stabilization techniques such as fine-grained credit assignment and trajectory filtering, instability remains pervasive and often leads to training collapse. We argue that this instability stems from inefficient exploration in multi-turn settings, where policies continue to generate low-information actions that neither reduce uncertainty nor advance task progress. To address this issue, we propose Token- and Turn-level Policy Optimization (T$^2$PO), an uncertainty-aware framework that explicitly controls exploration at fine-grained levels. At the token level, T$^2$PO monitors uncertainty dynamics and triggers a thinking intervention once the marginal uncertainty change falls below a threshold. At the turn level, T$^2$PO identifies interactions with negligible exploration progress and dynamically resamples such turns to avoid wasted rollouts. We evaluate T$^2$PO in diverse environments, including WebShop, ALFWorld, and Search QA, demonstrating substantial gains in training stability and performance improvements with better exploration efficiency. Code is available at: this https URL.
Comments: 25 pages, 7 figures, 8 tables. Accepted to ICML 2026 as a Spotlight Paper
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.02178 [cs.AI]
  (or arXiv:2605.02178v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2605.02178

arXiv-issued DOI via DataCite

Submission history

From: Haixin Wang [view email]
[v1] Mon, 4 May 2026 03:15:56 UTC (4,774 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org