跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分46

EasyPPO:稳定 critic 是 PPO 训练 LLM 的关键

AI 导读

EasyPPO 通过仅对 actor 做超长过滤、按采样回报标准差倒数加权 critic 回归、并适度缩小 critic mini-batch,解决了 PPO 中 critic 的两类失稳问题。

正文

View PDF HTML (experimental)

Abstract:A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.36802 [cs.LG]
  (or arXiv:2609.36802v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.36802

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Qiuyang Mang [view email]
[v1] Tue, 29 Sep 2026 06:15:00 UTC (785 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org