HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-05-27精选AI 评分71
DenoiseRL:通过恢复嘈杂前缀来引导推理模型
AI 导读
DenoiseRL是一种强化学习框架,旨在提升大语言模型的推理能力。它无需依赖更强的教师模型或精心筛选的困难数据集,而是通过在弱模型产生的失败推理轨迹上进行基于恢复的优化来直接学习,将错误转化为改进机会。这种方法提供了更丰富多样的学习信号,提升了探索效率。实验表明,DenoiseRL在竞争性的数学和通用推理基准测试中,持续优于强在策略RL基线,并能随着训练难度增加促进更强的自我纠正行为。
推荐理由
做 RL for reasoning 的团队该看这篇,它把训练信号从“依赖强模型”转向“从弱模型的错误中学习”,可能降低对昂贵 teacher 的依赖,是个架构层面的新思路。
正文
Abstract:Reinforcement learning has become a central paradigm for advancing reasoning in large language models, yet most existing methods still depend on stronger teacher models or heavily curated difficult datasets, limiting scalable capability improvement. In this paper, we introduce DenoiseRL, a reinforcement learning framework that substitutes external supervision with recovery-oriented optimization over failures from weak models. Instead of relying on stronger supervision or carefully engineered data, DenoiseRL learns directly from noisy reasoning prefixes by converting them into opportunities for improvement, while exercising fine-grained control over the noise intensity, making training more scalable and effective. This yields a richer and more diverse learning signal, improving exploration efficiency by leveraging imperfect model behavior. Empirically, DenoiseRL consistently outperforms strong on-policy RL baselines across competitive mathematical reasoning tasks and interactive decision-making tasks, while promoting stronger self-corrective behavior as training difficulty increases, highlighting an effective and scalable pathway for improving agentic reasoning capabilities of large language models.
| Comments: | 17 pages, 5 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2605.28421 [cs.AI] |
| (or arXiv:2605.28421v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2605.28421 arXiv-issued DOI via DataCite |
Submission history
From: Caijun Xu [view email]
[v1]
Wed, 27 May 2026 12:52:58 UTC (678 KB)
[v2]
Thu, 30 Jul 2026 08:43:29 UTC (550 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org