HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分34
SECRET:用源条件化中继引导缓解音视频大语言模型的幻觉
AI 导读
针对音视频大语言模型(AVLLM)的"来源混淆接地幻觉",研究者通过路径干预与表征分析发现"问题中继"机制:问题状态会携带干扰模态的线索,削弱对所需模态证据的接地。
正文
Abstract:Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confused grounding hallucination}$, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a $\textbf{question-relay}$ mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose $\textbf{SECRET}$ ($\textbf{S}$ourc$\textbf{E}$-$\textbf{C}$onditioned $\textbf{RE}$lay s$\textbf{T}$eering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.37568 [cs.CL] |
| (or arXiv:2609.37568v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37568 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yu Zhang [view email]
[v1]
Tue, 29 Sep 2026 13:37:31 UTC (862 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org