Representative voice cloning and zero/few-shot TTS models
| Name | Model/approach | Cloning type | Evaluation metrics (as reported) | Year |
|---|---|---|---|---|
| Tacotron 2 + WaveNet (Shen et al., 2018) | Seq2Seq acoustic model + autoregressive vocoder | Fine-tuning (speaker adaptation) | MOS, speaker similarity (Sim) | 2018 |
| d-vector TTS (Jia et al., 2018) | Speaker embeddings (d-vector) + Tacotron 2 | Few-shot | MOS, speaker similarity (Sim) | 2018 |
| x-vector TTS (Wan et al., 2018) | x-vector embeddings + neural vocoder | Few-shot | MOS, speaker similarity (Sim) | 2018 |
| AdaSpeech (Chen et al., 2021) | Adaptive TTS with meta-learning | Few-shot | MOS, Mel Cepstral Distortion (MCD) | 2021 |
| Meta-StyleSpeech (Min et al., 2021) | Meta-learning + style adaptation | Few-shot | MOS, speaker similarity (Sim) | 2021 |
| YourTTS (Casanova et al., 2022) | VITS-based, multilingual + adversarial training | Zero-shot | MOS, speaker similarity (Sim) | 2022 |
| ProDiff (Huang et al., 2022) | Progressive fast diffusion model | Multi-speaker TTS | MOS, real-time factor (RTF) | 2022 |
| DiffGAN-TTS (Liu et al., 2022) | Diffusion + adversarial training | Multi-speaker TTS | MOS, speaker similarity (Sim) | 2022 |
| VALL-E (Wang et al., 2023) | Neural codec language model | Zero-shot | MOS, speaker similarity (Sim) | 2023 |
| NaturalSpeech 2 (Shen et al., 2023) | Diffusion-based TTS | Zero-shot | MOS, CER/WER, speaker similarity (Sim) | 2023 |
| UnitSpeech (Kim et al., 2023a) | Discrete units + diffusion modeling | Few-shot/zero-shot | MOS, speaker similarity (Sim) | 2023 |
| XTTS (Casanova et al., 2023) | VQ-VAE + GPT + diffusion decoder | Zero-shot | MOS, speaker similarity (Sim) | 2023 |
| ElevenLabs (2023) | Commercial neural TTS system | Zero-shot | MOS (reported) | 2023 |
| NaturalSpeech 3 (Ju et al., 2024) | Factorized codec + diffusion modeling | Zero-shot | MOS, WER, speaker similarity (Sim) | 2024 |
| VALL-E 2 (Chen et al., 2024) | Improved codec-based transformer (successor of VALL-E) | Zero-shot | MOS, speaker similarity (Sim) | 2024 |
| Name | Model/approach | Cloning type | Evaluation metrics (as reported) | Year |
|---|---|---|---|---|
| Tacotron 2 + WaveNet ( | Seq2Seq acoustic model + autoregressive vocoder | Fine-tuning (speaker adaptation) | MOS, speaker similarity (Sim) | 2018 |
| d-vector | Speaker embeddings (d-vector) + Tacotron 2 | Few-shot | MOS, speaker similarity (Sim) | 2018 |
| x-vector | x-vector embeddings + neural vocoder | Few-shot | MOS, speaker similarity (Sim) | 2018 |
| AdaSpeech ( | Adaptive | Few-shot | MOS, Mel Cepstral Distortion ( | 2021 |
| Meta-StyleSpeech ( | Meta-learning + style adaptation | Few-shot | MOS, speaker similarity (Sim) | 2021 |
| YourTTS ( | VITS-based, multilingual + adversarial training | Zero-shot | MOS, speaker similarity (Sim) | 2022 |
| ProDiff ( | Progressive fast diffusion model | Multi-speaker | MOS, real-time factor ( | 2022 |
| DiffGAN-TTS ( | Diffusion + adversarial training | Multi-speaker | MOS, speaker similarity (Sim) | 2022 |
| VALL-E ( | Neural codec language model | Zero-shot | MOS, speaker similarity (Sim) | 2023 |
| NaturalSpeech 2 ( | Diffusion-based | Zero-shot | MOS, CER/WER, speaker similarity (Sim) | 2023 |
| UnitSpeech ( | Discrete units + diffusion modeling | Few-shot/zero-shot | MOS, speaker similarity (Sim) | 2023 |
| VQ-VAE + | Zero-shot | MOS, speaker similarity (Sim) | 2023 | |
| Commercial neural | Zero-shot | 2023 | ||
| NaturalSpeech 3 ( | Factorized codec + diffusion modeling | Zero-shot | MOS, WER, speaker similarity (Sim) | 2024 |
| VALL-E 2 ( | Improved codec-based transformer (successor of VALL-E) | Zero-shot | MOS, speaker similarity (Sim) | 2024 |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.