跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分38

Imagine3D-LLM:让多模态大模型先"想象"3D 场景再作答

AI 导读

Imagine3D-LLM 让多模态大模型在图像 token 后追加一组可学习的 summary token,将其解码为紧凑的 3D Gaussian Splatting 表示,并以光度重建损失监督、与标准下一 token 预测联合训练。仅 summary token 接受重建监督,却能增强 LLM 底层图像特征的跨帧对应,在多个空间推理与 3D 理解基准上持续优于此前方法。

正文

Authors:Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong

View PDF HTML (experimental)

Abstract:Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.
Comments: NeurIPS 2026; Project Page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as: arXiv:2609.38177 [cs.CV]
  (or arXiv:2609.38177v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.38177

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jaewoo Jung [view email]
[v1] Tue, 29 Sep 2026 17:59:52 UTC (6,756 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org