跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分43

FocusVTC:自适应分辨率的高效视觉文本压缩

AI 导读

FocusVTC 通过自适应分辨率打破视觉文本压缩的固定分辨率权衡,在保留通用多模态能力的同时压缩输入长度。在 RULER v1 上以 72 DPI 取得 87.4 分、2.9 倍输入压缩,超过 Glyph 的 57.5 分;LongBench 得分 56.40 高于文本输入基线的 55.86,MRCR 宏平均提升 13.91 分,在线端到端延迟较纯文本提速 2.79 倍。

正文

View PDF HTML (experimental)

Abstract:Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
Comments: 23 pages, 10 figures. Code: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.36651 [cs.CV]
  (or arXiv:2609.36651v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.36651

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Fangzhi Zhong [view email]
[v1] Tue, 29 Sep 2026 03:51:06 UTC (23,022 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org