Table 3.

Mean Absolute Difference of Hierarchical ED: differences between the predicted and the ground-truth hierarchical ED values. The column Longer “Longer Segments” denotes the longer segments used to predict shorter segments; “GT” indicates that ground-truth values were employed. For example, under the “GT” condition, we used the ground-truth utterance-level ED to predict the word-level ED, whereas under the “Predicted” condition, we utilized the predicted utterance-level ED

Hierarchical ED ConditionHierarchical ED Difference
TTS ModelPred ModeLonger SegmentsPhonemesWordsUtteranceAvg.
ExternalSingle-Step0.13330.12830.05940.1070
ExternalMulti-StepPredicted0.13450.12970.05870.1077
ExternalMulti-StepGT0.12140.12810.05870.1028
VASingle-Step0.13580.12940.05990.1084
VA(Multi-Step)Multi-StepPredicted0.13560.12980.06010.1085
VA(Multi-Step)Multi-StepGT0.12300.12720.06010.1034

or Create an Account

Close subscription notice
Close access options