HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分42
Just MLPs:为多模态语言模型高效重建视觉状态
AI 导读
研究提出 δ-Vision,用轻量低秩适配器构建逐层视觉记忆,替代视觉 token 在 Transformer 中的重复演化,同时保留全部视觉 token 供文本检索。该方法发现视觉到文本的信息流集中在低维子空间,且轻量 MLP 能以高余弦相似度和低重建误差逼近各层视觉状态。在图像和视频基准上,δ-Vision 在相当或更低计算量下准确率超过视觉 token 剪枝基线,且不丢弃视觉 token。
正文
Abstract:Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose $\delta$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, $\delta$-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.
| Comments: | 21 pages, 5 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.34972 [cs.CV] |
| (or arXiv:2609.34972v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34972 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jingdi Lei [view email]
[v1]
Mon, 28 Sep 2026 11:52:45 UTC (399 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org