Table 4

Performance of conventional and proposed PFG models (prosodic factor generator and prominence generator) on the preprocessed BC2013 dataset. The conventional model utilized text, while the proposed PFG model utilized text and emotion soft labels as input. The L2 loss of utterance-level prosodic factors and word-level prominence were calculated, respectively.

ModelInputL2 loss
Prosodic factor generatorText [26]0.062
 Text+EmoSoftLabel0.054
Prominence generatorText [34]0.023
 Text+EmoSoftLabel0.018

or Create an Account

Close subscription notice
Close access options