HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-05-04精选AI 评分70
T^2PO:面向稳定多轮智能体强化学习的不确定性引导探索控制框架
AI 导读
多轮强化学习训练常因探索效率低下而不稳定。为此,研究团队提出T^2PO框架,在细粒度层面实施不确定性引导的探索控制。在令牌级别,它监测不确定性动态,当边际变化低于阈值时触发思考干预;在轮次级别,它识别探索进展可忽略的交互并动态重采样,以避免无效计算。在WebShop、ALFWorld和Search QA等多个环境中的评估表明,T^2PO显著提升了训练稳定性与任务性能,并实现了更高效的探索。相关代码已开源。
推荐理由
多轮 agent 训练最怕训着训着崩了,这篇从 token 和 turn 两级控制探索的思路很妙,直接把低效 rollout 砍掉,稳定性和效率都上去了,做 RLHF 或 agent RL 的可以认真看一下。
正文
Abstract:Recent progress in multi-turn reinforcement learning (RL) has significantly improved reasoning LLMs' performances on complex interactive tasks. Despite advances in stabilization techniques such as fine-grained credit assignment and trajectory filtering, instability remains pervasive and often leads to training collapse. We argue that this instability stems from inefficient exploration in multi-turn settings, where policies continue to generate low-information actions that neither reduce uncertainty nor advance task progress. To address this issue, we propose Token- and Turn-level Policy Optimization (T$^2$PO), an uncertainty-aware framework that explicitly controls exploration at fine-grained levels. At the token level, T$^2$PO monitors uncertainty dynamics and triggers a thinking intervention once the marginal uncertainty change falls below a threshold. At the turn level, T$^2$PO identifies interactions with negligible exploration progress and dynamically resamples such turns to avoid wasted rollouts. We evaluate T$^2$PO in diverse environments, including WebShop, ALFWorld, and Search QA, demonstrating substantial gains in training stability and performance improvements with better exploration efficiency. Code is available at: this https URL.
| Comments: | 25 pages, 7 figures, 8 tables. Accepted to ICML 2026 as a Spotlight Paper |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2605.02178 [cs.AI] |
| (or arXiv:2605.02178v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2605.02178 arXiv-issued DOI via DataCite |
Submission history
From: Haixin Wang [view email]
[v1]
Mon, 4 May 2026 03:15:56 UTC (4,774 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org