HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分38
SCOPD:面向高效视觉语言模型的稀疏上下文同策略自蒸馏
AI 导读
针对推理型视觉语言模型(VLM)视觉 token 序列过长导致推理昂贵的问题,研究者提出稀疏上下文同策略自蒸馏框架 SCOPD:学生模型基于剪枝后的视觉 token 生成推理轨迹,由拥有完整上下文的教师模型对同一同策略前缀进行监督,无需真实答案、架构改动或额外推理开销。
正文
Authors:Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki Karaimer, Mozhgan Nasr Azadani, Jiayun Luo, Minh Ngoc Le, Gholamali Aminian, Hugo Buurmeijer, Yongchao Chen, Leonid Sigal, Igor Gilitschenski, Konstantinos G. Derpanis, Marco Pavone, Babak Taati
Abstract:Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.34044 [cs.CV] |
| (or arXiv:2609.34044v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34044 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ahmadreza Jeddi [view email]
[v1]
Mon, 28 Sep 2026 00:09:07 UTC (5,663 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org