Audio Flamingo Next:面向语音、声音与音乐的下一代开源音频-语言模型
Audio Flamingo 系列发布下一代模型 AF-Next,支持长达 30 分钟的复杂音频输入,并引入 Temporal Audio Chain-of-Thought 技术,将中间推理步骤显式对应到时间戳。该模型基于超过 100 万小时的新增数据进行课程式训练,在 20 个音频理解与推理基准测试中显著超越同规模开源模型,部分指标超过更大规模的闭源模型。团队同步开源了 AF-Next-Instruct、AF-Next-Think 和 AF-Next-Captioner 三个变体。
NVIDIA 开源的这个音频大模型在 20 多个基准上大幅超越前代,长音频理解尤其凶猛,做语音、音乐和声音应用的值得把论文翻一遍。
Authors:Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Zhifeng Kong, Siddharth Gururani, Sang-gil Lee, Jaehyeon Kim, Aya Aljafari, Chao-Han Huck Yang, Sungwon Kim, Ramani Duraiswami, Dinesh Manocha, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping
Abstract:We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3, AF-Next introduces: (i) a stronger foundational audio-language model that significantly improves accuracy across diverse audio understanding tasks; (ii) scalable strategies for constructing large-scale audio understanding and reasoning data beyond existing academic benchmarks; (iii) support for long and complex audio inputs up to 30 minutes; and (iv) Temporal Audio Chain-of-Thought, a new reasoning paradigm that explicitly grounds intermediate reasoning steps to timestamps in long audio, enabling fine-grained temporal alignment and improved interpretability. To enable these capabilities, we first conduct a systematic analysis of Audio Flamingo 3 to identify key gaps in audio understanding and reasoning. We then curate and scale new large-scale datasets totaling over 1 million hours to address these limitations and expand the existing AudioSkills-XL, LongAudio-XL, AF-Think and AF-Chat datasets. AF-Next is trained using a curriculum-based strategy spanning pre-training, mid-training and post-training stages. Extensive experiments across 20 audio understanding and reasoning benchmarks, including challenging long-audio tasks, show that AF-Next outperforms similarly sized open models by large margins and remains highly competitive with and sometimes surpasses, much larger open-weight and closed models. Beyond benchmark performance, AF-Next exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. In addition to all data, code and methods, we open-source 3 variants of AF-Next, including AF-Next-Instruct, AF-Next-Think and AF-Next-Captioner.
| Comments: | Project website: this https URL |
| Subjects: | Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2604.10905 [cs.SD] |
| (or arXiv:2604.10905v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2604.10905 arXiv-issued DOI via DataCite |
Submission history
From: Sreyan Ghosh [view email]
[v1]
Mon, 13 Apr 2026 02:11:56 UTC (7,410 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org