HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分41
Soft Spatial Reasoning:用自适应软思考缓解 LVLM 空间推理的过早离散化
AI 导读
研究者提出 Soft Spatial Reasoning,一个面向 LVLM 空间任务的训练后框架,在每步推理中混合 token 嵌入形成连续软状态,而非硬性选定单个 token。
正文
Abstract:Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at this https URL
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.38717 [cs.CV] |
| (or arXiv:2609.38717v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38717 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dongxiao Zhu [view email]
[v1]
Wed, 30 Sep 2026 00:51:48 UTC (6,539 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org