Figure 2.
Three T T S architectures show Grad T T S and D V T, V I T S, and Tacotron 2 pipelines with text encoder, alignment, decoder, and vocoder stages.The diagram shows three labelled architectures a Grad T T S and D V T, b V I T S, and c Tacotron 2. In a, text encoder connects to alignment, then to speech decoder with noise input, producing latent speech, followed by a vocoder generating a waveform. In b, the text encoder connects to alignment, then to a speech decoder with stacked flow blocks and noise input, followed by a vocoder producing a waveform. In c, the text encoder connects to soft attention, then to the speech decoder, producing speech representation, followed by the vocoder generating waveform, with a feedback connection from the output to attention.

Simplified diagrams of the evaluated TTS architectures. (a) Diffusion-based models (GradTTS, DVT), (b) The flow-based model (VITS), (c) The autoregressive-based baseline (tacotron 2). previously presented inFeng et al. (2024) 

or Create an Account

Close subscription notice
Close access options