The workflow begins with language-independent vocal mimicry from a person speaking into a microphone. The vocal signal enters a Transformer and then a neural vocoder. A newly constructed paired dataset contains Vocal Mimicry 1, Vocal Mimicry 2 through Vocal Mimicry N paired with Sound Effect 1, Sound Effect 2 through Sound Effect N. The paired dataset connects to the Transformer and neural vocoder. The neural vocoder produces a synthesised sound effect represented by an explosion symbol with a waveform.Schematic image of PronounSE. The method synthesizes sound effects from vocal mimicry using Transformer (Vaswani et al., 2017) and a neural vocoder
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.