HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分43
FocusVTC:自适应分辨率的高效视觉文本压缩
AI 导读
FocusVTC 通过自适应分辨率打破视觉文本压缩的固定分辨率权衡,在保留通用多模态能力的同时压缩输入长度。在 RULER v1 上以 72 DPI 取得 87.4 分、2.9 倍输入压缩,超过 Glyph 的 57.5 分;LongBench 得分 56.40 高于文本输入基线的 55.86,MRCR 宏平均提升 13.91 分,在线端到端延迟较纯文本提速 2.79 倍。
正文
Abstract:Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
| Comments: | 23 pages, 10 figures. Code: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.36651 [cs.CV] |
| (or arXiv:2609.36651v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36651 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fangzhi Zhong [view email]
[v1]
Tue, 29 Sep 2026 03:51:06 UTC (23,022 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org