跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分49

距离去掉视觉编码器还有多远?无编码器多模态预训练的缩放定律

AI 导读

研究对比了无编码器与有编码器多模态大模型的缩放定律,发现去掉视觉编码器会把多模态目标的算力最优分配推向更大模型,文本目标则几乎不变。两者在文本目标上损失-算力前沿几乎重合,但多模态目标上无编码器模型小规模时落后,预计约 10^22 FLOPs 时追平。无编码器时语言模型通过视觉专属适配接管其角色:视觉 token 双向交互随算力增长愈发有益,视觉处理前移,视觉 token 的专家路由更集中。

正文

View PDF HTML (experimental)

Abstract:Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.35457 [cs.CV]
  (or arXiv:2609.35457v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.35457

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lin Chen [view email]
[v1] Mon, 28 Sep 2026 15:48:41 UTC (1,342 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org