HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分41
Think Before You Score:面向视觉生成的思想奖励模型 TRM
AI 导读
研究者提出 Thinking Reward Model(TRM),先为每个样本生成自适应评分标准再做评估,输出细粒度 pointwise 奖励,并引入 PD-GRPO 缓解成对偏好优化导致的分数极化。在图像生成与编辑奖励建模基准上,TRM 在开源奖励模型中达到 SOTA,并可与闭源方案竞争;作为强化学习奖励信号还能稳定提升多种视觉生成模型。
正文
Authors:Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
Abstract:Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.
| Comments: | 31 pages |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.37372 [cs.CV] |
| (or arXiv:2609.37372v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37372 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yang Shi [view email]
[v1]
Tue, 29 Sep 2026 12:28:00 UTC (5,481 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org