跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分41

预训练 LLM 中的 Softmax 近似:模型敏感性与内核加速

AI 导读

研究在十个冻结的 decoder-only 模型(0.5B-72B)上近似推理期 softmax,发现可大幅削减分配概率的位置数与行内分辨率,但对相同位置做均匀加权会带来损害,且固定分辨率预算的分配位置与预算大小同样重要。

正文

View PDF HTML (experimental)

Abstract:On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.33586 [cs.LG]
  (or arXiv:2609.33586v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.33586

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shangzhen Zhu [view email]
[v1] Sun, 27 Sep 2026 14:03:42 UTC (585 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org