HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分36
在注意力机制中用量化 Softmax 预训练 Transformer
AI 导读
研究在注意力机制中用 K-interval attention 量化 softmax 对 Transformer 预训练的影响,该方法用 K+1 个网格值近似指数函数。
正文
Abstract:Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, and compare these choices in pretraining experiments matched on model, data, and optimizer. Detaching the row extrema leaves the forward computation unchanged but produces a delayed increase in validation loss. With hard rounding at K=4, min-max calibration and a pre-normalization surrogate incur a large loss gap; changing either choice substantially reduces it. At 124M parameters and 2.5B training tokens, fixed-window calibration with a post-normalization surrogate yields a validation loss gap of +0.019 nats relative to softmax at K=4, and with a pre-normalization surrogate yields +0.004 nats at K=16.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.33591 [cs.LG] |
| (or arXiv:2609.33591v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33591 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shangzhen Zhu [view email]
[v1]
Sun, 27 Sep 2026 14:09:52 UTC (701 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org