HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-08精选AI 评分88
Agon:基于隐式对抗评分的跨模型竞争强化学习框架
AI 导读
Agon 让两个模型互为评分者,通过竞争性强化学习提升推理能力。在 DeepMath 困难子集上,基于 Qwen3 的 Agon 将 GRPO 的 pass@1 翻倍,增益约为未训练的 Mixture-of-Agents 的 8 倍。该结果在 Qwen3.5、Gemma 4 等模型族及编程代码任务上得到复现,推理时采用两阶段级联:一个模型起草,另一个阅读后作答。
推荐理由
这篇论文用「让两个模型互搏」代替了只看最终答案的RL训练,直接解决了推理模型「写得多、想得少」的毛病。0.6B小模型用这方法能超过4B的GRPO模型,做推理训练的人应该认真读一读。
正文
Abstract:Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other's graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen its work, so reasoning is judged implicitly during training, with no process labels and no reward model. Because both models are optimized, each faces a progressively stronger rival, which single-model RL cannot provide. The two need only be comparably strong and behaviorally different. At inference the pair deploys as it trains, a two-stage cascade in which one model drafts and the other answers after reading the draft. On the hard split of DeepMath with Qwen3, this doubles GRPO's pass@1, roughly eight times the gain of an untrained Mixture-of-Agents pass over the same base. The ordering replicates on competitive-programming code and across model families (Qwen3.5, Gemma 4). For now the models talk in text; the next step is to let them reason together in latent space.
| Comments: | 15 pages, 7 figures, 8 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.07690 [cs.LG] |
| (or arXiv:2607.07690v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.07690 arXiv-issued DOI via DataCite |
Submission history
From: Vladislav Beliaev [view email]
[v1]
Wed, 8 Jul 2026 17:49:14 UTC (25 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org