跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-31精选AI 评分80

Zero-Mem:为 LLM 智能体实现零 token 记忆操作

AI 导读

Zero-Mem 提出零 token 记忆操作,除最终问答外,记忆构建、检索与校准等所有环节均不调用 LLM 或消耗其输入输出 token,仅用编码器计算。该方法保留原始交互轨迹,通过实体-上下文图与时间层级双视图组织证据,在长记忆与长上下文问答基准上达到竞争性表现,并将记忆操作时间成本较最快基线降低 57.6%。代码将在同行评审后公开。

推荐理由

Zero-Mem将记忆操作与LLM完全解耦,通过实体图和时序层级互补检索,在LoCoMo和HotpotQA上全面超越生成式记忆基线,效率提升明显,对Agent记忆系统设计有直接借鉴意义。

正文

Authors:Yilin Xiao, Zhehan Zhu, Yujing Zhang, Jin Chen, Zijin Hong, Luyao Zhuang, Qinggang Zhang, Shengyuan Chen, Xiaocao Ouyang, Lingfei Ren, Xiao Huang

View PDF HTML (experimental)

Abstract:LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and mediating their retrieval adds recurring token and time costs, while omitted or merged details can obscure the original evidence. We ask whether structured memory access requires generation at all. Zero-Mem introduces \emph{zero-token memory operations}: no step outside final question answering invokes an LLM or consumes LLM input or output tokens; encoder computation is accounted for separately. Zero-Mem preserves original interaction traces as its source of record. It organizes the traces in two complementary ways. An entity--context graph exposes connections across interactions, while a temporal hierarchy preserves conversational locality and session state. For each query, Zero-Mem weighs the two views, retrieves from both, and follows their structure to recover supporting relations or surrounding context. Deterministic calibration first discards conflicting evidence and then keeps the reader's answer grounded in the retrieved traces. Only the final-QA reader invokes an LLM. Across long-memory and long-context question-answering benchmarks, Zero-Mem achieves competitive performance while eliminating LLM calls and LLM-token consumption from memory operations. With the same final-QA reader and context budget, it reduces memory-operation time cost by 57.6\% relative to the fastest compared baseline. Ablations support the contribution of the two views and their query-dependent coordination. Overall, the results show that structured agent memory need not generate an intermediate representation of the past. After peer review, the code and implementation details will be available at \textcolor{blue}{this https URL}.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2607.29377 [cs.CL]
  (or arXiv:2607.29377v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.29377

arXiv-issued DOI via DataCite

Submission history

From: Yilin Xiao [view email]
[v1] Fri, 31 Jul 2026 13:01:06 UTC (414 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org