HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分38
EmoRES:面向情感语音生成的残差增强向量引导方法
AI 导读
EmoRES 提出一种免训练的情感 TTS 向量引导方法,将情感向量拆分为共享分量与残差分量并分别控制,无需重训主干模型。
正文
Abstract:Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.
| Comments: | Work done at Meta. Code at this https URL |
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2609.38157 [cs.SD] |
| (or arXiv:2609.38157v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38157 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kuan-Po Huang [view email]
[v1]
Tue, 29 Sep 2026 17:59:05 UTC (363 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org