跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分33

APM-Bench:面向第一视角流式视频助手的跨会话持久记忆基准

AI 导读

APM-Bench 将真实流式交互重构为多会话生活轨迹,包含 549 个会话、104 条轨迹和 2,719 个候选问题,覆盖客观与开放式问答。模型需借助持久记忆回答过往会话问题并主动响应,同时保持实时交互。评测显示现有方法在效用、延迟与存储之间存在明显权衡,难以同时实现可靠的长期召回、低开销和有效的跨会话主动协助。

正文

Authors:Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu, Jianhang Li, Zhimian Wen, Mingda Li, Shule Lu, Zhicheng Wang, Yuhan Guo, Xin Jin, Wenjun Zeng

View PDF HTML (experimental)

Abstract:To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.
Comments: 33 pages, 11 figures, 15 tables
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.37559 [cs.CV]
  (or arXiv:2609.37559v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.37559

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jianguo Huang [view email]
[v1] Tue, 29 Sep 2026 13:32:06 UTC (7,710 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org