The three labelled stages a, b, and c. Stage a shows clean speech combined with environmental noise under S N R conditions, negative 10 to 0, negative 15 to 0, and negative 20 to 0, producing noisy speech. This passes through A S R to produce a noisy transcription. Stage b shows T T S systems with multiple models and transcription conditions, including clean, W E R 070, and W E R 162, used to train a T T S model that outputs clean speech. Stage c shows input text from test data entering T T S models in inference mode, followed by evaluation models and outputs labelled experimental evaluation of synthetic speech and experimental analysis.Overview of the experimental framework. (a) Noisy transcriptions are simulated by adding environmental noise at various SNR conditions to clean speech, which is then transcribed using an ASR system, (b) TTS models are trained on clean speech paired with different transcription conditions (clean, WER-070, WER-162, etc.), (c) Trained models are evaluated through perceptual listening tests and analytical analysis using input text from both clean and noisy test sets
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.