Figure 1.
A workflow maps language-independent vocal mimicry through a Transformer and neural vocoder to synthesise sound effects using a paired vocal mimicry and sound-effect dataset.The workflow begins with language-independent vocal mimicry from a person speaking into a microphone. The vocal signal enters a Transformer and then a neural vocoder. A newly constructed paired dataset contains Vocal Mimicry 1, Vocal Mimicry 2 through Vocal Mimicry N paired with Sound Effect 1, Sound Effect 2 through Sound Effect N. The paired dataset connects to the Transformer and neural vocoder. The neural vocoder produces a synthesised sound effect represented by an explosion symbol with a waveform.

Schematic image of PronounSE. The method synthesizes sound effects from vocal mimicry using Transformer (Vaswani et al., 2017) and a neural vocoder

or Create an Account

Close Modal
Close Modal