跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分39

UMM-Reflection:用交错强化学习让统一多模态模型学会原生反思

AI 导读

UMM-Reflection 通过交错强化学习,让统一多模态模型在一条轨迹内完成"诊断—修改—再观察"的完整反思循环,轨迹级优势同时更新反思 token 与基于流的图像修改,推理时无需外部验证器。

正文

View PDF HTML (experimental)

Abstract:Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.35767 [cs.CV]
  (or arXiv:2609.35767v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.35767

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yijia Fan [view email]
[v1] Mon, 28 Sep 2026 17:59:36 UTC (20,973 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org