跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分38

重新思考潜在视觉推理:将潜在推理锚定在视觉证据上

AI 导读

针对潜在视觉推理中潜在 token 对改变正确答案的图像扰动响应微弱这一"潜在证据信用缺口",研究者提出 ReaLVR,用正确与模型生成错误答案的对比确定监督位置,用匹配与不匹配的视觉证据指定应保留的内容。在三个模型家族上 ReaLVR 均优于现有 LVR 基线,在 Qwen2.5-VL-7B 上取得五项任务平均 63.7% 的最高成绩,并首次将潜在空间视觉推理扩展至 235B 规模。

正文

Authors:Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu, Jiaye Wu, Zheng Zhang, Xiang Xu, Xuanbai Chen, Farhan Tejani, Jakub Zablocki, Julia Xu, Yifan Xing

View PDF HTML (experimental)

Abstract:Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
Comments: 39 pages. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.34563 [cs.CV]
  (or arXiv:2609.34563v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.34563

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xi Xiao [view email]
[v1] Mon, 28 Sep 2026 08:16:14 UTC (12,469 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org