Figure 1.
A three-part system diagram shows text to speech processing with emotion prediction, emotion control, and emotion editing.A three-part system diagram presents workflows for speech synthesis and emotion handling. The left section shows text to speech with emotion prediction, where input text is processed by a proposed system to predict hierarchical emotion descriptors from the text, which are then used to generate synthesised audio. The middle section shows text to speech with emotion control, where input text is processed to predict utterance level, word level, and phoneme level emotion descriptors, with an emotion control component influencing these predictions before synthesised audio is produced. The right section shows emotion editing, where input audio and its corresponding text are processed by the proposed system to extract hierarchical emotion descriptors from the audio, which are then modified through an emotion editing process to generate synthesised audio with adjusted emotional characteristics.

Inference diagram of the proposed system: (a) TTS with emotion prediction; (b) TTS with emotion control and (c) emotion editing. The hierarchical emotion distribution (ED) can be obtained in three ways: (1) directly predicted from the input text (“Emotion Prediction”), (2) predicted from the input text with user modifications (“Emotion Control”), or (3) extracted from the input audio and manually adjusted by users (“Emotion Editing”)

or Create an Account

Close subscription notice
Close access options