跳到正文
原文
X:Rohan Paul (@rohanpaul_ai)· X:Rohan Paul (@rohanpaul_ai)·· 6 天前AI 评分60

StepFun 发布 StepAudio 3 Music 音乐生成基础模型

AI 导读

StepFun 与 ACE Studio 发布 StepAudio 3 Music,可将文本描述和歌词生成完整歌曲,支持歌曲生成、器乐生成、音乐翻唱、人声编排四种工作流,并基于 ABC-COT 先规划音乐结构与编排再合成。

正文

StepFun has released StepAudio 3 Music, a model that turns a text description and lyrics into a finished song.

The interesting engineering result is that better audio reconstruction did not necessarily produce better music generation.

The technical report compares single-codebook and residual vector quantization approaches. The final tokenizer uses one token stream at 50 Hz with 65,536 entries, followed by a flow-matching DiT renderer.

Why this matters: the representation has to preserve sound AND give the autoregressive model a sequence it can predict reliably. Optimizing the codec in isolation can miss that trade-off.

(In comment you will find some audio examples to give this architecture some context)

来源:X:Rohan Paul (@rohanpaul_ai) · x.com