HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分34
大语言模型注意力机制演进综述:机制、权衡与新兴趋势
AI 导读
一篇综述将 LLM 注意力机制的发展统一为"模型内部上下文记忆"视角,提出涵盖记忆表示、更新、访问、读取与整合的五维分析框架。研究基于 14 个主要模型谱系、59 条发布级记录和 11 个高性能开放权重端点,重建了机制层面的演进与架构采用情况。综述指出,高效序列架构设计正从孤立的 Attention 算子转向上下文记忆的组织、生命周期与选择性使用,并据此提出有状态多维记忆路由假设。
正文
Abstract:Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update, Access, Readout, and Integration---describing what is represented, how it changes, what is query-eligible, how it is read, and how readouts form outputs. This lens compares overlapping research lines without imposing one computational model.
We reconstruct mechanism-level developments and architectural adoption using 59 release-level records from 14 major model lineages and 11 high-performing open-weight endpoints. First, explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions. Second, heterogeneous architectures increasingly coordinate across network depth: layer-wise composition distributes complementary memory processing across representational stages, while cross-layer reuse carries selected memory and routing artifacts forward. Depth thus becomes a dimension along which contextual memory is constructed and managed. Third, these developments motivate a stateful multidimensional memory-routing hypothesis: persistent memory is organized across temporal scope, network depth, substrate type, and representation granularity, while coordinated Sparse Write and Sparse Read determine what is maintained and what contributes to each query. Overall, efficient sequence architecture design increasingly concerns the organization, lifecycle, and selective use of contextual memory rather than an isolated Attention operator.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.39661 [cs.CL] |
| (or arXiv:2609.39661v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.39661 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhentao Tan [view email]
[v1]
Wed, 30 Sep 2026 13:03:01 UTC (1,714 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org