HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分41
预训练 LLM 中的 Softmax 近似:模型敏感性与内核加速
AI 导读
研究在十个冻结的 decoder-only 模型(0.5B-72B)上近似推理期 softmax,发现可大幅削减分配概率的位置数与行内分辨率,但对相同位置做均匀加权会带来损害,且固定分辨率预算的分配位置与预算大小同样重要。
正文
Abstract:On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.33586 [cs.LG] |
| (or arXiv:2609.33586v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33586 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shangzhen Zhu [view email]
[v1]
Sun, 27 Sep 2026 14:03:42 UTC (585 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org