跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分35

CorrGRPO:面向多奖励学习的相关性归一化 GRPO

AI 导读

针对 GRPO 在多奖励场景下归一化易被大尺度相关奖励主导的问题,研究者提出 CorrGRPO,将成对协方差归一化为 Pearson 相关系数,在保持中心化总奖励不变的同时平衡不同尺度奖励的影响。在代码生成、工具调用和智能体安全三类任务上,使用 0.5B 至 8B 参数模型与 GRPO 及其他变体对比,三个领域均取得提升,代码已开源。

正文

View PDF HTML (experimental)

Abstract:Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at this https URL.
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2609.36820 [cs.LG]
  (or arXiv:2609.36820v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.36820

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wenbin Hu [view email]
[v1] Tue, 29 Sep 2026 06:30:02 UTC (1,408 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org