HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分41
Pruned CTC:面向大词表 ASR 训练的内存高效方案
AI 导读
Pruned CTC 将 CTC 对齐计算限制在目标 token 与 blank 构成的小子集内,同时保留全词表归一化,并证明其损失与梯度同全词表 CTC 完全等价。
正文
Authors:Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
Abstract:Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
| Subjects: | Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2609.33645 [eess.AS] |
| (or arXiv:2609.33645v2 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33645 arXiv-issued DOI via DataCite |
Submission history
From: Yifan Yang [view email]
[v1]
Sun, 27 Sep 2026 15:10:59 UTC (558 KB)
[v2]
Tue, 29 Sep 2026 16:59:49 UTC (558 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org