HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分38
VoxPolyMem:面向多说话人语音对话的交互感知多模态记忆框架
AI 导读
VoxPolyMem 是一个面向多说话人语音对话的交互感知多模态记忆框架,结合增量说话人识别与交互记忆、事实记忆、参与者画像三层记忆结构,并在 VoxPolyBench 上取得 85.0 总分,超出最强基线 23.6 分。
正文
Abstract:Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom. To this end, we propose VoxPolyMem, an interaction-aware multimodal memory framework combining incremental speaker identification with a memory hierarchy comprising interaction memory, fact memory, and participant profiles. We formulate retrieval as sequential decision-making, where an agent rewrites queries and selects retrieval tools and memory layers based on accumulated evidence to address information gaps. We further introduce Evidence-Gain GRPO (EG-GRPO), which uses round-wise credit assignment to encourage complementary evidence acquisition. We also construct VoxPolyBench to evaluate memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution in multi-party spoken conversations. VoxPolyMem achieves an overall score of 85.0 on VoxPolyBench, surpassing the strongest evaluated baseline by 23.6 points. On Mem-Gallery and H2HMem-Multi, it scores 89.6 and 74.4, respectively, exceeding the strongest evaluated public memory baselines by over 8 points each. These results highlight its potential for persistent, personalized assistance in multi-party multimodal interactions. Code and datasets are available at this https URL
| Subjects: | Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Sound (cs.SD); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2609.32522 [cs.AI] |
| (or arXiv:2609.32522v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.32522 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wenxu Jia [view email]
[v1]
Sat, 26 Sep 2026 12:04:49 UTC (3,845 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org