跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分38

VoxPolyMem:面向多说话人语音对话的交互感知多模态记忆框架

AI 导读

VoxPolyMem 是一个面向多说话人语音对话的交互感知多模态记忆框架,结合增量说话人识别与交互记忆、事实记忆、参与者画像三层记忆结构,并在 VoxPolyBench 上取得 85.0 总分,超出最强基线 23.6 分。

正文

View PDF HTML (experimental)

Abstract:Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom. To this end, we propose VoxPolyMem, an interaction-aware multimodal memory framework combining incremental speaker identification with a memory hierarchy comprising interaction memory, fact memory, and participant profiles. We formulate retrieval as sequential decision-making, where an agent rewrites queries and selects retrieval tools and memory layers based on accumulated evidence to address information gaps. We further introduce Evidence-Gain GRPO (EG-GRPO), which uses round-wise credit assignment to encourage complementary evidence acquisition. We also construct VoxPolyBench to evaluate memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution in multi-party spoken conversations. VoxPolyMem achieves an overall score of 85.0 on VoxPolyBench, surpassing the strongest evaluated baseline by 23.6 points. On Mem-Gallery and H2HMem-Multi, it scores 89.6 and 74.4, respectively, exceeding the strongest evaluated public memory baselines by over 8 points each. These results highlight its potential for persistent, personalized assistance in multi-party multimodal interactions. Code and datasets are available at this https URL
Subjects: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2609.32522 [cs.AI]
  (or arXiv:2609.32522v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.32522

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wenxu Jia [view email]
[v1] Sat, 26 Sep 2026 12:04:49 UTC (3,845 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org