跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分42

Tacit-TTS:从自回归解码到掩码预测的高效无转写语音克隆

AI 导读

Tacit-TTS 是蒸馏自 IndexTTS2 的无转写零样本语音克隆系统,用掩码非自回归生成替代自回归文本到语义解码,并引入免训练声学长度的估计与 ReFlow 蒸馏加速流匹配渲染器。在 2 个英语和 2 个普通话数据集上,其对超过 5 秒的语句生成速度比 IndexTTS2 快 10 倍以上,且无需参考语音转写,可支持跨语言及婴儿咿呀、合成乱语等非词汇参考。

正文

View PDF HTML (experimental)

Abstract:TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
Comments: Under Review
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as: arXiv:2609.38658 [eess.AS]
  (or arXiv:2609.38658v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2609.38658

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jian Chen [view email]
[v1] Tue, 29 Sep 2026 23:25:21 UTC (1,701 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org