HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分40
AgSpec:把基于检索的投机解码推向编码智能体流水线极限
AI 导读
AgSpec 是一个面向编码智能体流水线的检索式投机解码框架,从会话、工作区和全局语料库中检索草稿 token,并按智能体离线画像设定草稿长度上限、再依据验证反馈在线调整。在两个仓库级多智能体编码基准上,AgSpec 在多数设置下优于五种检索式草稿器和 EAGLE-3,batch size 1 时生成吞吐较自回归解码最高提升 4.37 倍,batch size 16 时最高 4.76 倍。
正文
Abstract:Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37$\times$ at batch size 1 and 4.76$\times$ at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.01108 [cs.CL] |
| (or arXiv:2610.01108v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01108 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sukmin Cho [view email]
[v1]
Thu, 1 Oct 2026 05:48:29 UTC (728 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org