Performance of conventional and proposed PFG models (prosodic factor generator and prominence generator) on the preprocessed BC2013 dataset. The conventional model utilized text, while the proposed PFG model utilized text and emotion soft labels as input. The L2 loss of utterance-level prosodic factors and word-level prominence were calculated, respectively.
| Model | Input | L2 loss |
|---|---|---|
| Prosodic factor generator | Text [26] | 0.062 |
| Text+EmoSoftLabel | 0.054 | |
| Prominence generator | Text [34] | 0.023 |
| Text+EmoSoftLabel | 0.018 |
| Model | Input | L2 loss |
|---|---|---|
| Prosodic factor generator | Text [ | 0.062 |
| Prominence generator | Text [ | 0.023 |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.