Table 1.

Speech Quality Test Results: MUSHRA naturalness scores with 95% confidence interval and Word Error Rate (WER). The column “GT or Pred” indicates whether we use ground-truth hierarchical ED (“GT”) or a text-predicted version. In the “TTS Model” column, “VA” and “VA(Multi-Step)” denote the TTS models employing the Single-Step hierarchical ED Variance Adaptor (Figure 4(d)) and the Multi-Step hierarchical ED Variance Adaptor (Figure 4(b)), respectively. Finally, the “Pred Mode” column specifies whether we predict the hierarchical ED sequentially from longer to shorter segments (Multi-Step) or in parallel for all segments (Single-Step)

Hierarchical EDSpeech Quality
GT or PredTTS ModelPred ModeMUSHRA ()WER ()
— Ground-Truth Speech Samples —79.4± 1.92.16
GTExternal61.6± 2.23.37
GTVA57.5± 2.63.11
GTVA(Multi-Step)62.2± 2.32.48
PredictedExternalSingle-Step50.7± 2.43.80
PredictedExternalMulti-Step54.0± 2.33.25
PredictedVASingle-Step52.2± 2.64.61
PredictedVA(Multi-Step)Multi-Step53.2± 2.42.45

or Create an Account

Close subscription notice
Close access options