Figure 3.
A multi-part system diagram shows hierarchical emotion distribution extraction, prediction, training, and emotion editing in a speech synthesis model.A multi-part system diagram presents the architecture and workflows of a speech synthesis model using hierarchical emotion distributions. The upper left section shows an overall model where input audio is processed by a hierarchical emotion distribution extractor to produce hierarchical emotion embeddings, which are combined with text processed through grapheme to phoneme conversion, text encoding, a variance adaptor, and a decoder to generate a mel spectrogram. The upper right section shows the training process for a hierarchical emotion distribution predictor, where text is processed through grapheme to phoneme conversion and a text encoder, followed by utterance level, word level, and phoneme level emotion distribution predictors, with training optimised by minimising differences from ground truth hierarchical emotion distributions. The lower left section shows an emotion editing process, where predicted utterance level and phoneme level emotion distributions are modified through emotion control to produce edited emotion distributions, which are then embedded and combined with the text processing pipeline to generate an edited mel spectrogram. The lower right section shows an example of emotion editing, where emotion intensity values for angry, happy, sad, and surprise are adjusted, including an increase in the sad intensity for the final word.

Training and Inference Diagrams of the proposed framework using external integration (“External”): (a) Overall diagram; (b) Training diagram of hierarchical emotion distribution (ED) predictor; (c) Emotion editing (inference) diagram and (d) Example of emotion editing

or Create an Account

Close subscription notice
Close access options