HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分40
Where-OPD:用合成场景对 MLLM 做空间引导的在线自蒸馏
AI 导读
Where-OPD 提出一种面向多模态大模型(MLLM)的在线自蒸馏方法,让教师模型获得文本形式的空间定位引导,从而定位并整合多个相关图像区域的证据,学生模型仅凭图像和问题学习复现该行为。
正文
Abstract:On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: this https URL
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02117 [cs.CV] |
| (or arXiv:2610.02117v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02117 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sophia Sirko-Galouchenko [view email]
[v1]
Thu, 1 Oct 2026 17:34:10 UTC (2,413 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org