The diagram shows three labelled architectures a Grad T T S and D V T, b V I T S, and c Tacotron 2. In a, text encoder connects to alignment, then to speech decoder with noise input, producing latent speech, followed by a vocoder generating a waveform. In b, the text encoder connects to alignment, then to a speech decoder with stacked flow blocks and noise input, followed by a vocoder producing a waveform. In c, the text encoder connects to soft attention, then to the speech decoder, producing speech representation, followed by the vocoder generating waveform, with a feedback connection from the output to attention.Simplified diagrams of the evaluated TTS architectures. (a) Diffusion-based models (GradTTS, DVT), (b) The flow-based model (VITS), (c) The autoregressive-based baseline (tacotron 2). previously presented inFeng et al. (2024)
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.