跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分38

VGGT-Diff:融合视觉几何与扩散模型的稀疏视角新视角合成

AI 导读

VGGT-Diff 将 VGGT-Ω 的视觉几何 latent 路由进预训练视频扩散模型,用于稀疏视角新视角合成。每个视觉 token 关联 3D 点和置信度,经置信度感知的 Visual Geometry Router(VGR)转为查询对齐的 latent 条件,Point-Track Residual Consistency(PTRC)沿可靠 3D 轨迹正则化预测残差。

正文

Authors:Kangjie Chen, Xiangyu Li, Dongbin Zhang, Chaoda Zheng, Shijia Chen, Jinhao Deng, Hongbin Lin, Choo Sin Wai, Minqi Wang, Minghao Yang, Dake Zhong, Guorui Song, Yu Zhang, Xianming Liu, Boyang Wang

View PDF HTML (experimental)

Abstract:We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-{\Omega} into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at this https URL.
Comments: Project page: this https URL, Code at: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.33253 [cs.CV]
  (or arXiv:2609.33253v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.33253

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kangjie Chen [view email]
[v1] Sun, 27 Sep 2026 05:57:25 UTC (20,032 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org