HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分33
FactorEngram:面向语言模型的基级门控分解式 N-gram 记忆
AI 导读
FactorEngram 提出一种分解式 n-gram 记忆机制,用共享基向量字典上的稀疏系数替代传统单体嵌入,并让上下文对每个记忆分量单独门控。在 340M 和 1B 参数 Transformer 骨干上,该方法提升了语言建模与下游任务表现;消融实验显示,将记忆分支插入中间层的注意力子层之前是有效配置。
正文
Abstract:Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.35578 [cs.CL] |
| (or arXiv:2609.35578v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35578 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jingbo Zhou [view email]
[v1]
Mon, 28 Sep 2026 16:35:18 UTC (536 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org