Table 2.

Emotion Expressiveness Test Results with 95% confidence interval: MUSHRA similarity scores, Mel-Cepstral Distortion (MCD), Pitch/Energy Distortion (Pitch/Energy), and Frame Disturbance (FD). The column “GT or Pred” indicates whether we use ground-truth hierarchical ED (“GT”) or a text-predicted version. In the “TTS Model” column, “VA” and “VA(Multi-Step)” denote the TTS models employing the Single-Step hierarchical ED Variance Adaptor (Figure 4(d)) and the Multi-Step hierarchical ED Variance Adaptor (Figure 4(b)), respectively. Finally, the “Pred Mode” column specifies whether we predict the hierarchical ED progressively from longer to shorter segments (Multi-Step) or in parallel for all segments (Single-Step)

Hierarchical EDEmotion Expressiveness
GT or PredTTS ModelPred ModeMUSHRA ()MCD ()Pitch ()Energy ()FD ()
GTExternal61.9± 2.15.88± 0.1015.6± 1.00.363± 0.02224.8± 3.4
GTVA55.8± 2.66.48± 0.2316.1± 1.10.386± 0.02325.5± 3.6
GTVA(Multi-Step)61.9± 2.15.62± 0.1115.5± 1.10.348± 0.02022.4± 2.9
PredictedExternalSingle-Step47.2± 2.37.59± 0.1418.2± 1.10.438± 0.02746.3± 6.5
PredictedExternalMulti-Step51.9± 2.26.89± 0.1216.9± 1.20.409± 0.02442.6± 4.7
PredictedVASingle-Step48.2± 2.57.23± 0.2016.7± 1.00.426± 0.02541.4± 4.7
PredictedVA(Multi-Step)Multi-Step49.1± 2.26.91± 0.1217.2± 1.10.416± 0.02546.3± 5.1

or Create an Account

Close subscription notice
Close access options