跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-11精选AI 评分70

GigaChat Audio:支持120分钟输入的时间感知音频大语言模型

AI 导读

GigaChat Audio 是一种时间感知的音频大语言模型,支持长达120分钟的输入,并能生成带有明确时间戳的问答、片段描述和摘要。该模型通过将周期性时间标记与连续音频 token 交错,并利用级联合成数据管道进行大规模训练,在短时和长时基准上均实现了强时间定位精度。研究团队已开源模型权重及超过1万小时的时序数据集。

推荐理由

首个大规模时间感知音频LLM开源,把长音频问答的时间定位做到可验证,且附带完整数据集和消融分析,做语音产品的该认真看一下。

正文

View PDF HTML (experimental)

Abstract:Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at this https URL.
Comments: Accepted to Interspeech 2026. Model and dataset: this https URL
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
Cite as: arXiv:2607.10387 [eess.AS]
  (or arXiv:2607.10387v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2607.10387

arXiv-issued DOI via DataCite

Submission history

From: Georgii Gospodinov [view email]
[v1] Sat, 11 Jul 2026 16:34:25 UTC (617 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org