HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分37
GAGAR:面向代码智能体 RL 的分组智能体评分与优势重分配
AI 导读
GAGAR 是一个面向代码智能体强化学习的质量感知信用重分配框架,通过 SFT 训练的智能体评分器在同一工作区内联合检查并排序通过测试的轨迹,对低排名轨迹降权并按比例重缩放优势值以保持总和不变。
正文
Abstract:Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.32577 [cs.CL] |
| (or arXiv:2609.32577v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.32577 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jinhao Dong [view email]
[v1]
Sat, 26 Sep 2026 12:50:04 UTC (397 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org