跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-04-13精选AI 评分77

RationalRewards:推理奖励在训练与测试时扩展视觉生成

AI 导读

研究团队发布8B参数奖励模型RationalRewards,基于PARROT框架从偏好数据恢复高质量推理依据,实现打分前生成多维度显式批评。该模型仅用同类基线1/10到1/20训练数据即达开源SOTA水平,性能比肩Gemini-2.5-Pro。其结构化推理奖励不仅为强化学习提供细粒度训练信号,更在测试时通过生成-批评-优化循环无需参数更新即可改进输出,在多个基准上匹敌甚至超越传统RL微调效果。

推荐理由

这篇论文把视觉生成的奖励模型从黑盒标量变成了带推理的结构化批评,而且发现测试时的生成-批评-改进循环能匹敌昂贵的RL微调,做图像生成和编辑的人应该认真读一下。

正文

View PDF HTML (experimental)

Abstract:Most reward models for visual generation reduce rich human judgments to a single unexplained score, discarding the reasoning that underlies preference. We show that teaching reward models to produce explicit, multi-dimensional critiques before scoring transforms them from passive evaluators into active optimization tools, improving generators in two complementary ways: at training time, structured rationales provide interpretable, fine-grained rewards for reinforcement learning; at test time, a Generate-Critique-Refine loop turns critiques into targeted prompt revisions that improve outputs without any parameter updates. To train such a reward model without costly rationale annotations, we introduce Preference-Anchored Rationalization (PARROT), a principled framework that recovers high-quality rationales from readily available preference data through anchored generation, consistency filtering, and distillation. The resulting model, RationalRewards (8B), achieves state-of-the-art preference prediction among open-source reward models, competitive with Gemini-2.5-Pro, while using 10-20x less training data than comparable baselines. As an RL reward, it consistently improves text-to-image and image-editing generators beyond scalar alternatives. Most strikingly, its test-time critique-and-refine loop matches or exceeds RL-based fine-tuning on several benchmarks, suggesting that structured reasoning can unlock latent capabilities in existing generators that suboptimal prompts fail to elicit.
Comments: Project Page: this https URL ; Code, Dataset, Models are released
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2604.11626 [cs.AI]
  (or arXiv:2604.11626v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2604.11626

arXiv-issued DOI via DataCite

Submission history

From: Haozhe Wang [view email]
[v1] Mon, 13 Apr 2026 15:38:09 UTC (11,526 KB)
[v2] Tue, 14 Apr 2026 06:06:15 UTC (11,527 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org