HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分42
Tacit-TTS:从自回归解码到掩码预测的高效无转写语音克隆
AI 导读
Tacit-TTS 是蒸馏自 IndexTTS2 的无转写零样本语音克隆系统,用掩码非自回归生成替代自回归文本到语义解码,并引入免训练声学长度的估计与 ReFlow 蒸馏加速流匹配渲染器。在 2 个英语和 2 个普通话数据集上,其对超过 5 秒的语句生成速度比 IndexTTS2 快 10 倍以上,且无需参考语音转写,可支持跨语言及婴儿咿呀、合成乱语等非词汇参考。
正文
Abstract:TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
| Comments: | Under Review |
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD) |
| Cite as: | arXiv:2609.38658 [eess.AS] |
| (or arXiv:2609.38658v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38658 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jian Chen [view email]
[v1]
Tue, 29 Sep 2026 23:25:21 UTC (1,701 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org