跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分35

WUSH-KV:用数据自适应变换实现 KV Cache 量化

AI 导读

WUSH-KV 是一种面向低比特 KV cache 的量化方法,利用矩阵乘积两个因子的二阶统计量构建数据感知变换,为 key 和 value 分别构造变换,其中 value 变换折叠进模型权重、key 变换在 RoPE 之后应用。

正文

View PDF HTML (experimental)

Abstract:KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.38121 [cs.LG]
  (or arXiv:2609.38121v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38121

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jiale Chen [view email]
[v1] Tue, 29 Sep 2026 17:50:50 UTC (314 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org