HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分32
DARA:面向多奖励强化学习的密度感知奖励聚合方法
AI 导读
针对多奖励强化学习中各目标学习进度不均的问题,研究者提出密度感知奖励聚合方法 DARA,通过逆平方根密度校正,为激活频率较低的奖励信号赋予更高权重,且不改变底层策略优化目标。在工具调用任务上,DARA 达到高格式合规所需训练步数比 GDPO 最多减少 26%;在数学推理任务上,接近饱和的长度合规所需步数最多减少 65%,最终性能保持竞争力。
正文
Abstract:Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at this https URL.
| Comments: | 24 pages, 8 figures |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.00574 [cs.LG] |
| (or arXiv:2610.00574v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00574 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haotian Zhai [view email]
[v1]
Wed, 30 Sep 2026 18:45:26 UTC (441 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org