Mean Absolute Difference of Hierarchical ED: differences between the predicted and the ground-truth hierarchical ED values. The column Longer “Longer Segments” denotes the longer segments used to predict shorter segments; “GT” indicates that ground-truth values were employed. For example, under the “GT” condition, we used the ground-truth utterance-level ED to predict the word-level ED, whereas under the “Predicted” condition, we utilized the predicted utterance-level ED
| Hierarchical ED Condition | Hierarchical ED Difference | |||||
|---|---|---|---|---|---|---|
| TTS Model | Pred Mode | Longer Segments | Phonemes | Words | Utterance | Avg. |
| External | Single-Step | – | 0.1333 | 0.1283 | 0.0594 | 0.1070 |
| External | Multi-Step | Predicted | 0.1345 | 0.1297 | 0.0587 | 0.1077 |
| External | Multi-Step | GT | 0.1214 | 0.1281 | 0.0587 | 0.1028 |
| VA | Single-Step | – | 0.1358 | 0.1294 | 0.0599 | 0.1084 |
| VA(Multi-Step) | Multi-Step | Predicted | 0.1356 | 0.1298 | 0.0601 | 0.1085 |
| VA(Multi-Step) | Multi-Step | GT | 0.1230 | 0.1272 | 0.0601 | 0.1034 |
| Hierarchical | Hierarchical | |||||
|---|---|---|---|---|---|---|
| Pred Mode | Longer Segments | Phonemes | Words | Utterance | Avg. | |
| External | Single-Step | – | 0.1333 | 0.1283 | 0.0594 | 0.1070 |
| External | Multi-Step | Predicted | 0.1345 | 0.1297 | 0.1077 | |
| External | Multi-Step | 0.0587 | ||||
| Single-Step | – | 0.1358 | 0.1294 | 0.1084 | ||
| VA(Multi-Step) | Multi-Step | Predicted | 0.1356 | 0.1298 | 0.0601 | 0.1085 |
| VA(Multi-Step) | Multi-Step | 0.0601 | ||||
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.