HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分41
WaveFront Decoding:面向循环语言模型的无训练自推测解码框架
AI 导读
WaveFront Decoding(WFD)是一种面向循环语言模型的无训练自推测解码框架,利用中间循环输出作为草稿预测、并借助权重共享批量处理不同位置与循环深度的 token 状态。
正文
Abstract:Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore concurrently batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B. The code is available at this https URL.
| Comments: | 18 pages, 7 figures, 8 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.23033 [cs.LG] |
| (or arXiv:2609.23033v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.23033 arXiv-issued DOI via DataCite |
Submission history
From: Hyeongju Ha [view email]
[v1]
Sat, 19 Sep 2026 14:08:07 UTC (475 KB)
[v2]
Sat, 26 Sep 2026 09:59:28 UTC (499 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org