Figure 3.
A two-part workflow trains a Transformer on paired vocal mimicry and sound-effect mel spectrograms, then converts vocal mimicry into a synthesised waveform.The upper section presents the training phase of the Transformer. Vocal mimicry mel spectrogram X passes through a Pre net and Embedding X. Positional encoding is added before a 3-layer Transformer Encoder. Sound-effect mel spectrogram Y passes through a Pre net and Embedding Y. Positional encoding is added before a 3-layer Transformer Decoder, which also receives the encoder output. The decoder output passes through a linear layer to produce predicted mel spectrogram Y hat. A Post net produces Y hat post net. The loss is the L 1 loss between Y hat and Y plus the L 1 loss between Y hat post net and Y. The lower section presents the synthesis process. A vocal mimicry waveform passes through S T F T plus Mel F B preprocessing to produce a vocal mimicry mel spectrogram. The trained Transformer converts this into a converted mel spectrogram. Hi Fi G A N reconstructs the synthesised waveform.

Architecture of PronounSE. This figure illustrates the inference-time synthesis flow of PronounSE and the training process of Transformer, which is trained to map Mel-spectrograms of vocal mimicry to those of corresponding sound effects

or Create an Account

Close Modal
Close Modal