Figure 1.
A three-stage pipeline shows noisy transcription simulation, model training, and inference evaluation using A S R, T T S, and evaluation models.The three labelled stages a, b, and c. Stage a shows clean speech combined with environmental noise under S N R conditions, negative 10 to 0, negative 15 to 0, and negative 20 to 0, producing noisy speech. This passes through A S R to produce a noisy transcription. Stage b shows T T S systems with multiple models and transcription conditions, including clean, W E R 070, and W E R 162, used to train a T T S model that outputs clean speech. Stage c shows input text from test data entering T T S models in inference mode, followed by evaluation models and outputs labelled experimental evaluation of synthetic speech and experimental analysis.

Overview of the experimental framework. (a) Noisy transcriptions are simulated by adding environmental noise at various SNR conditions to clean speech, which is then transcribed using an ASR system, (b) TTS models are trained on clean speech paired with different transcription conditions (clean, WER-070, WER-162, etc.), (c) Trained models are evaluated through perceptual listening tests and analytical analysis using input text from both clean and noisy test sets

or Create an Account

Close subscription notice
Close access options