跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分34

自注意力在竞争下的检索容量

AI 导读

研究通过保留每个注意力头、层和查询中注意力权重最高的 token 并测量 NLL 增幅,估算语言模型维持损失容差所需的"有效注意力集合"大小。基于注意力的选择显著优于随机选择,扩展上下文会增大所需集合大小,而重归一化保留权重可大幅降低该规模。

正文

View PDF HTML (experimental)

Abstract:How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size needed to stay within a chosen loss tolerance. Relatively small selected sets can keep NLL close to the full-attention baseline, although the required size varies across models. Attention-based selection substantially outperforms random selection. Selected sets exhibit geometric structure, although geometric separation alone does not establish that model loss is preserved. Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range. Experiments with a fixed supporting fact show that additional background pushes its tokens down the attention ranking and reduces their attention mass. Renormalizing the retained weights can substantially reduce the required set size, showing that it also depends on how selected representations are combined. Conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. These results provide a way to measure effective attention set size in language models and investigate its dependence on context, competition, and aggregation.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.37879 [cs.CL]
  (or arXiv:2609.37879v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.37879

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Timur Mudarisov [view email]
[v1] Tue, 29 Sep 2026 15:53:04 UTC (2,174 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org