Figure 4.
A multi-part system diagram shows variance adaptor designs for hierarchical emotion distribution modelling and emotion control in speech synthesis.A multi-part system diagram presents different variance adaptor designs used in a speech synthesis model with hierarchical emotion distributions. The upper left section shows the overall model workflow, where text from input audio is processed through grapheme to phoneme conversion, a text encoder, a multi-step hierarchical emotion distribution variance adaptor, and a decoder to produce a predicted mel spectrogram. The upper middle section details the multi-step hierarchical emotion distribution variance adaptor, where utterance level, word level, and phoneme level emotion distribution predictors operate sequentially and feed into a variance adaptor. The upper right section shows a standard variance adaptor that predicts duration, pitch, and energy using dedicated predictors, alongside a one-step hierarchical emotion distribution variance adaptor that predicts hierarchical emotion distributions in a single stage. The lower section shows an emotion editing workflow, where predicted utterance level and phoneme level emotion distributions are modified through an emotion control mechanism to create edited emotion distributions, which are then passed through the variance adaptor and decoder to generate edited speech output. An example of emotion control demonstrates changes in emotion intensity values, including an increase in the sad intensity for the final word.

Training and Inference Diagrams of the proposed framework using variance adaptor integration (“VA”): (a) Overall diagram; (b) Diagram of sequential hierarchical emotion distribution (hierarchical ED) variance adaptor; (c) Diagram of variance adaptor; (d) Diagram of parallel hierarchical ED variance adaptor (e) Emotion editing (inference) diagram; (f) Example of emotion editing and (g) Example of emotion control

or Create an Account

Close subscription notice
Close access options