Figure 4.
Spectrograms compare a reference with Vocal Mimicry, T Foley, Ours and Stable Audio 2.0 outputs across six speakers.A reference spectrogram appears beside a grid of spectrograms for speakers m 01, m 02, m 03, m 04, m 05 and f 01. The grid contains four rows labelled Vocal Mimicry, T Foley, Ours and Stable Audio 2.0. Frequency axes extend from 0 to 8192 hertz. The reference extends to approximately 4 seconds. Vocal Mimicry spectrograms contain varied frequency patterns across the six speakers. T Foley spectrograms also vary across speakers, with several containing concentrated activity near the beginning. Ours contains broad frequency content near the beginning that narrows towards lower frequencies over time across all six speakers. Stable Audio 2.0 displays a similar narrowing pattern across the six speakers.

Examples of the reference sound, vocal mimicry from six speakers and sounds synthesized using PronounSE, T-Foley (Chung et al., 2024) and Stable Audio 2.0 (Evans et al., 2024). The synthesized results are shown in the order of T-Foley, PronounSE and Stable Audio 2.0

or Create an Account

Close Modal
Close Modal