跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分37

LoopVL:循环式视觉智能

AI 导读

LoopVL 将 Loop Transformer 扩展到视觉语言模型,通过 Module-Loop 与 Model-Loop 组合,用共享模块迭代更新统一的视觉语言状态,并从零开始完成语言预训练、多模态训练与后训练。它在多模态理解与视觉推理基准上超过多款同规模及更大规模的非循环模型,还出现跨循环视觉注意力显著转移的 Visual Aha Moments。

正文

Authors:Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li, Zhonghua Wang, Fei Luo, Mingxuan Wang, Xue Yang, Shiwei liu, Yanbiao Ma, Junchi Yan, Jungong Han

View PDF HTML (experimental)

Abstract:We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as: arXiv:2609.38426 [cs.CV]
  (or arXiv:2609.38426v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.38426

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhe Qian [view email]
[v1] Tue, 29 Sep 2026 19:17:25 UTC (30,955 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org