Table 2.

Intelligibility evaluation results (CER and WER in %) on speech synthesized from models trained under seven different transcription conditions

Training conditionsCleanWER-070WER-162WER-242WER-358WER-437WER-548
Models / (%)CERWERCERWERCERWERCERWERCERWERCERWERCERWER
Tacotron 23.37.66.110.97.113.111.219.511.522.517.728.727.745.3
Vits3.46.84.18.25.110.37.314.111.921.917.831.627.946.7
GradTTS3.47.03.77.64.08.64.38.75.210.66.914.08.916.5
Dvt3.36.93.67.44.48.94.99.35.410.96.512.87.413.9
Dvt23.26.83.37.44.18.44.28.64.49.65.610.86.913.2
Vits-dp3.17.03.67.64.18.65.910.98.215.212.322.019.333.6

or Create an Account

Close subscription notice
Close access options