跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分57

Galahad 让 LLM 重复读取文档成为一次性成本,在 llama.cpp 上将召回测试提升到 100/100

AI 导读

论文提出面向 vLLM、SGLang 和 llama.cpp 的内存层 Galahad,把模型对同一段文本的读取从重复计算变为一次性成本。

正文

View PDF HTML (experimental)

Abstract:A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and this http URL that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on this http URL at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG); Performance (cs.PF)
Cite as: arXiv:2609.39358 [cs.CL]
  (or arXiv:2609.39358v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.39358

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sietse Schelpe [view email]
[v1] Wed, 30 Sep 2026 09:16:40 UTC (28 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org