HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分42
WhiteMatter:通过 KV Source Mixing 实现全连接跨层连接
AI 导读
WhiteMatter 让 Transformer 每一层都能调用任意深度的历史 token 表示,通过学习型 mixer 为当前上下文挑选最有用的深度并合并进共享 KV cache 通道。
正文
Abstract:When generating text, a Transformer produces representations of past tokens at every layer, but each layer can normally use only representations from the same depth. This restriction prevents the model from fully reusing information it has already computed. We introduce WhiteMatter, which allows every layer to draw on past-token representations from any depth. A learned mixer selects the most useful depths for the current context and combines their representations into shared key-value (KV) cache channels. Sharing these channels across layers can reduce the cache size. Given the same number of training tokens, WhiteMatter with a full-size cache performs comparably to a standard Transformer with 50% more layers. With half the KV cache, WhiteMatter outperforms matched standard Transformers at two model scales, up to 1.3B parameters. Cross-layer connections, however, introduce dependencies that slow training and prompt processing. We address this problem with cyclic iteration, which updates interleaved groups of tokens in turn while processing the tokens within each group in parallel. On a reference model trained with exact autoregressive execution, cyclic iteration converges 12.5x faster than standard Jacobi iteration.
| Comments: | Code available at this https URL |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.7; I.2.6 |
| Cite as: | arXiv:2608.18486 [cs.CL] |
| (or arXiv:2608.18486v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.18486 arXiv-issued DOI via DataCite |
Submission history
From: Wenbo Zhang [view email]
[v1]
Wed, 19 Aug 2026 03:24:51 UTC (7,630 KB)
[v2]
Sun, 27 Sep 2026 18:24:11 UTC (7,662 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org