HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分43
LEAP:面向长音视频感知的学习式分块证据检索框架
AI 导读
LEAP 将长录音切分为固定时长区块,用轻量定位阶段为每个区块的候选窗口打分,再把排名最高的窗口汇聚到一次有界回答阶段重新编码,使回答输入与峰值上下文不随录音时长增长。
正文
Authors:Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, Liang Peng
Abstract:Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
| Comments: | 39 pages, 16 figures |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.39938 [cs.CL] |
| (or arXiv:2609.39938v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.39938 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Juyi Lin [view email]
[v1]
Wed, 30 Sep 2026 15:19:12 UTC (270 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org