Figure 5.
Five line charts show W E R percent across models from clean to 548 for Tacotron 2, V I T S D P, Grad T T S, D V T, and D V T 2 using clean and noisy style text.The five line charts are labelled Tacotron 2, V I T S D P, Grad T T S, D V T, and D V T 2. The horizontal axis lists models clean, 070, 162, 242, 358, 437, and 548. The vertical axis shows W E R percent from 0 to 60. Each chart has two lines for clean style text and noisy style text. In all models, values increase from clean to 548. Tacotron 2 rises from about 7 to about 45 for clean style and from about 2 to about 53 for noisy style. V I T S D P rises from about 7 to about 34 and from about 3 to about 49. Grad T T S rises from about 7 to about 17 and from about 4 to about 27. D V T rises from about 7 to about 14 and from about 2 to about 22. D V T 2 rises from about 7 to about 13 and from about 3 to about 25.

Intelligibility (WER %) results of the sensitivity test. A significant performance gap emerges between “clean-style text” (blue lines) and “noisy-style text” (orange lines) as the training data noise level increases, revealing the latent domain separation phenomenon

or Create an Account

Close subscription notice
Close access options