跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分40

PanoVLN:面向全景视觉语言导航的有效方案

AI 导读

PanoVLN 通过让模型从单张全景图预测更长动作序列,并引入置信度引导执行(CGE)策略动态决定执行步数,在 R2R-CE 和 RxR-CE Val-Unseen 上成功率分别超越此前 SOTA 11.9% 和 8.7%。

正文

View PDF HTML (experimental)

Abstract:Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward: more complete visual context should enable better-informed navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration. However, we find that simply replacing perspective images with panoramas yields only limited gains. Our diagnosis suggests that fully exploiting wider visibility requires modifications to action prediction, training supervision, and visual representation. First, wider visibility supports longer-horizon action planning. We make the model predict longer action sequences, enabling larger turns and subsequent movement from a single panorama. Specifically, we introduce a confidence-guided execution (CGE) strategy that dynamically determines how many predicted actions to execute before replanning. Second, wider visibility also brings more complex route choices. We therefore construct training routes with frequent branching points and clear instructions to provide targeted supervision for route selection. Third, panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks. We combine semantic and geometric features from RGB panoramas to capture both scene content and spatial layout without adding visual tokens. With a 4B backbone and RGB-only input, PanoVLN surpasses the previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen. Real-world experiments on a quadruped further demonstrate faster navigation with fewer pauses than prior VLN methods.
Comments: 22 pages, 14 figures. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
Cite as: arXiv:2609.34759 [cs.CV]
  (or arXiv:2609.34759v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.34759

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhen Wang [view email]
[v1] Mon, 28 Sep 2026 09:44:27 UTC (28,497 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org