HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分44
LT-OPD:用在线策略自蒸馏实现极端视觉 token 压缩
AI 导读
针对多模态大模型在极低视觉 token 预算下性能骤降的问题,研究者提出训练框架 LT-OPD,让仅保留少量视觉 token 的学生模型在自己生成的轨迹上接受冻结全 token 教师模型的分布监督,并引入逐步降低 token 预算的课程。
正文
Abstract:Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
| Comments: | Code is at this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.32353 [cs.CV] |
| (or arXiv:2609.32353v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.32353 arXiv-issued DOI via DataCite |
Submission history
From: Junxian Li [view email]
[v1]
Sat, 26 Sep 2026 08:16:07 UTC (2,817 KB)
[v2]
Thu, 1 Oct 2026 14:20:19 UTC (2,794 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org