跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分39

VoxMem:面向大型音频语言模型的多模态记忆基准

AI 导读

VoxMem 是面向大型音频语言模型(LALM)的多模态记忆基准,包含 3,196 个评测实例、34,743 段语音会话(177 小时),覆盖语音语义、说话人身份、副语言线索与环境声四类声学证据,以及信息抽取、多会话推理、时序追踪和拒答四类记忆操作,上下文预算从 8K 到 64K tokens。

正文

View PDF HTML (experimental)

Abstract:Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Comments: working in process
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as: arXiv:2609.32607 [eess.AS]
  (or arXiv:2609.32607v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2609.32607

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yang Xiao [view email]
[v1] Sat, 26 Sep 2026 13:27:12 UTC (28,079 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org