HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-17精选AI 评分73
NVIDIA 发布 Audio-Visual Flamingo:面向长视频的开放音频-视觉大语言模型
AI 导读
NVIDIA 推出 Nemotron-Labs-Audio-Visual Flamingo(AV-Flamingo),一个完全开源的音频-视觉大语言模型,专为理解和推理长视频中的音频、图像与复杂场景设计。
推荐理由
NVIDIA 这次把音频和视觉在长视频上的联合推理做成了全开源模型,AV-Skills 数据集和 TAVIT 推理框架是亮点,做视频理解的开发者可以直接拉下来用。
正文
Authors:Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping
Abstract:We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.
| Comments: | Project Page: this https URL |
| Subjects: | Audio and Speech Processing (eess.AS); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.16107 [eess.AS] |
| (or arXiv:2607.16107v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16107 arXiv-issued DOI via DataCite |
Submission history
From: Sreyan Ghosh [view email]
[v1]
Fri, 17 Jul 2026 16:41:15 UTC (8,913 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org