HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-11精选AI 评分70
GigaChat Audio:支持120分钟输入的时间感知音频大语言模型
AI 导读
GigaChat Audio 是一种时间感知的音频大语言模型,支持长达120分钟的输入,并能生成带有明确时间戳的问答、片段描述和摘要。该模型通过将周期性时间标记与连续音频 token 交错,并利用级联合成数据管道进行大规模训练,在短时和长时基准上均实现了强时间定位精度。研究团队已开源模型权重及超过1万小时的时序数据集。
推荐理由
首个大规模时间感知音频LLM开源,把长音频问答的时间定位做到可验证,且附带完整数据集和消融分析,做语音产品的该认真看一下。
正文
Abstract:Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at this https URL.
| Comments: | Accepted to Interspeech 2026. Model and dataset: this https URL |
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.10387 [eess.AS] |
| (or arXiv:2607.10387v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2607.10387 arXiv-issued DOI via DataCite |
Submission history
From: Georgii Gospodinov [view email]
[v1]
Sat, 11 Jul 2026 16:34:25 UTC (617 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org