HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分38
Imagine3D-LLM:让多模态大模型先"想象"3D 场景再作答
AI 导读
Imagine3D-LLM 让多模态大模型在图像 token 后追加一组可学习的 summary token,将其解码为紧凑的 3D Gaussian Splatting 表示,并以光度重建损失监督、与标准下一 token 预测联合训练。仅 summary token 接受重建监督,却能增强 LLM 底层图像特征的跨帧对应,在多个空间推理与 3D 理解基准上持续优于此前方法。
正文
Authors:Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
Abstract:Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.
| Comments: | NeurIPS 2026; Project Page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.38177 [cs.CV] |
| (or arXiv:2609.38177v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38177 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jaewoo Jung [view email]
[v1]
Tue, 29 Sep 2026 17:59:52 UTC (6,756 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org