Zyphra 发布开源实时语音克隆模型 ZONOS2
Zyphra 发布 Apache 2.0 许可的实时 TTS 模型 ZONOS2,采用 8B 总参数、900M 激活参数的 MoE 架构,主打高保真零样本语音克隆,权重已上线 Hugging Face。
原文给出架构、训练数据和评测细节,读者可以了解这个开源实时 TTS 模型在音色保真与生成稳定性之间的取舍。
Quantitative Comparison
Regarding quantitative metrics, first, we demonstrate that ZONOS2 excels on measures of naturalness and high-fidelity voice cloning, focusing especially on prosodic naturalness and speaker similarity.


In terms of model stability measured by word-error-rate (WER), ZONOS2 is comparable with other leading open-source and proprietary TTS systems. However, there is an interesting nuance to these comparisons. When we actually plot the reported scores of other models against the ground truth values, we observe that most TTS models score substantially lower than the ‘natural’ error rate of the evaluating ASR model vs the actual ground truth text. This is not necessarily overfitting but instead reflects the fact that many TTS systems are trained to produce speech that is substantially ‘cleaner’ and more intelligible to ASR systems than actual human speech.
This ultimately means that there is a trade-off between achieving strong WER metrics and the voice cloning faithfully matching the vocal characteristics of the source clip. With ZONOS2 we chose to emphasize vocal fidelity, even when the source clip contains background noise, unusual voices, or other distortions, vs always producing extremely clean ‘studio quality’ speech.
This effect can be visualized in more detail by plotting the multiple speaker metrics of the ground-truth audio vs other models. Here we can see many models often have a distinctly different ‘shape’ in the distribution of their outputs than the actual ground truth samples, implying that they subtly shift the distribution of their outputs during voice-cloning. By contrast, ZONOS2 sticks closely to the actual distribution within the ground-truth data resulting in more natural sounding audio.
However, despite our emphasis on voice-cloning fidelity, we appreciate there is also a strong case to be made for highly-stable and ‘studio quality’ speech, even if it requires departing somewhat from a truly faithful voice clone. For instance, a user may upload a poor quality voice-cloning sample with background noise, an unusual recording environment, or other audio distortions. In this case, they may want the TTS system to airbrush out these imperfections rather than faithfully replicate them in its generated audio. To serve both cases, we release ZONOS2 with two modes – a ‘stable’ mode which emphasizes clean output and an ‘expressive’ mode which focuses purely on voice-cloning naturalness and fidelity.
What's New in ZONOS2
ZONOS2 is a major step forward from Zonos-v0.1 Beta. It supports multilingual and code-switched audio generation up to one minute in length, uses simpler and more robust conditioning signals for improved stability and naturalness, and enables lifelike zero-shot voice cloning with a new ECAPA-TDNN speaker embedding model. With 20× the bandwidth of our previous speaker embedding model and new voice cloning training recipe, ZONOS2 captures significantly more vocal nuance and produces more convincing voice clones for a wide range of voices.

Schematic of the MoE block used in the ZONOS2 architecture. We follow our ZAYA MoE recipe including the novel MLP router we introduced for that model.
ZONOS2’s primary architectural novelty comes from adapting the powerful Mixture-of-Experts (MoE) transformer architecture to real-time TTS. Introducing MoE and removing dependence on Classifier Free Guidance (CFG) allows us to grow model size from 1.6B to 8B while improving real-time throughput by 4x compared to our previous model.
ZONOS2 predicts Descript Audio Codec (DAC) tokens, enabling it to generate studio-quality 44.1 kHz audio. DAC gives us access to exceptionally high-fidelity outputs, but modeling its tokens is significantly more challenging than working with simpler, lower-quality autoencoders. We overcome this complexity through increased model and data scale, allowing ZONOS2 to preserve audio quality while maintaining strong generation performance.

Animation of the forward pass of the ZONOS2 model. ZONOS2 autoregressively predicts DAC tokens that can then be decoded directly into 44 KHz audio. As in Zonos-v0.1, we used a delay pattern architecture to enable efficient parallel generation of DAC tokens while still properly maintaining serial dependencies across time.
Training and Architecture Details
Data Filtering
To improve on ZONOS-v0.1 Beta, we increased our training set from roughly 200,000 hours to over six million hours (or approximately 707 years) of audio. This expansion gives ZONOS2 substantially broader coverage of voices, languages, recording conditions, and text domains, while improving robustness to noise and atypical speaking patterns.
To navigate data at this scale, we introduce a new filtering method for web-scale audio data. In stage one we apply minimal inclusion criteria based on VAD and transcription validity. All audio that passes our VAD checks and receives a non-empty transcript is included in our training set. After this segmentation and labeling stage, we schedule increasingly strict inter-transcript agreement requirements over the course of training. This allows us to maximize data variety during pretraining, then progressively filter out unintelligible and low-quality audio during midtraining and annealing.
We find that scheduling transcript agreement across the three training stages lets us take advantage of the diversity of web-scale audio while avoiding many common TTS failure modes, including hallucinations, mispronunciations, repetitions, and other artifacts.
Training Recipe
ZONOS2 is trained on a simple autoregression task to predict a sequence of audio tokens given text and a speaker embedding extracted from source voice-clone audio. The output audio tokens are decoded back to raw speech waveforms for playback by the DAC decoder module. In order to model each of DAC’s ‘codebooks’ tokens, a step is conditioned on the prediction of lower order codebooks for that timestep re-ordered using a delay pattern such that the generation of each codebook for a given time step is generated sequentially.
Unlike ZONOS-v0.1 Beta, ZONOS2 represents text inputs as raw UTF-8 bytes, eliminating the need for an explicit phonemization step. This makes training more flexible, greatly expands the set of supported languages, and reduces errors caused by incorrect language labels and phonemization dictionaries. We find that byte tokenization is especially important for robust coverage of lower-resource languages, while also substantially improving performance on non-European languages such as Chinese, Korean, and Japanese, for which phonemization is substantially more challenging. Further, since it does not depend on explicit language tags, this approach natively enables code-switching in generated audio, as the model does not rely on hard coded per language tokenisation/phonemisation during inference, unlike the prior version.
ZONOS2 takes several conditioning signals that allow a user to control its generations. The model receives a speaker embedding as input which underpins its voice cloning capabilities. To enable flexible control of the generated speech, ZONOS2 also allows for the specification of target speaking-rate in eight available speeds. We support several quality dials for acoustic properties such as final estimated bandlimit for the generated audio as well as volume and estimated SNR.
ZONOS2 is trained in three phases. In the initial phase, the model is pre-trained over our entire dataset for 8 epochs without speaker cloning information and minimal transcript agreement filtering. This is followed by a shorter mid-training phase, where transcript agreement and sub-dataset selection are used to target clean high quality transcripts and audio in order to reduce hallucinations. Finally, both the speaker embedding, rate-of-speech conditioning, and quality conditioning are introduced in a short annealing training phase, using stricter filtering settings than the mid-training phase. The result of this three stage training scheme is reduced hallucinations with greater generalization to unseen voices and improved naturalness.
For more information, please see our ZONOS2 Hugging Face model card.
If you plan on self-hosting, ZONOS2 inference code is available on GitHub.
If you want to explore or use the model, it is available on Zyphra Cloud and as an API.
Introducing ZTTS1-Eval
Alongside our ZONOS2 model, we introduce a new TTS evaluation that we call ZTTS1-Eval. Throughout the evaluation process, it became clear that existing and widely-used TTS evals namely Seed-TTS-Eval and CV3-Eval have significant flaws.
Seed-TTS-Eval has three main problems. First, only Chinese and English are considered, which lacks linguistic diversity. Second, they rely on outdated models such as Whisper-Large for ASR. Thirdly, they only draw from Common Voice for English and DiDiSpeech for Chinese, both of which are read speech (with minimal prosodic variation) with a poor recording environment.
CV3-Eval partially improves upon Seed-TTS-Eval. Audio from seven additional languages are included from the FLEURS and Common Voice datasets, and tasks beyond accuracy and cloning (cross-lingual voice cloning, expressive voice continuation, emotion cloning, expressive voice cloning from web-crawled data) are added. However, the evaluation setup from seed-tts-eval is retained, using the same outdated models for ASR.
We set out to dramatically improve upon the existing state-of-the-art in TTS evaluations, and especially to produce a more robust, diverse, and less noisy evaluation than both Seed-TTS-Eval and CV3-Eval .
We release two evaluation sets. The first one consists of “Clean” audio from 9 languages using FLEURS-R, a restored version of the original FLEURS dataset, chosen for consistency and general acoustic quality. The second “In-The Wild” subsamples VoxBlink2 across 17 languages for acoustic and linguistic diversity. VoxBlink2 is composed of in-the-wild clips with highly variable quality, we choose samples across 10 uniform bins, spanning a 1-4 UTMOS quality score range to ensure fair representation across quality levels.
For evaluation we choose to swap older models for ones approaching the current state of the art. We use Qwen3-ASR for speech recognition (top of the Open ASR Leaderboard private set), ReDimNet for speaker similarity, MSR-UTMOS for acoustic quality.
To evaluate prosodic expressivity, we include the TTSDS Prosody metrics and Discretized Speech Weighted Edit Distance (DS-WED) in our eval. The TTSDS Prosody metrics measure distribution distances between generated and real speech, normalized by the real-to-noise distance. The set comprises Allosaurus SR (phoneme/s), HuBERT SR (semantic units/s), Pitch (mean F0), and MPM (masked prosody model embedding). DS-WED quantifies generation diversity by measuring the semantic unit distance between multiple generations of the same text.
来源:Zyphra Research(网页) · zyphra.com