Table 1.

FAD scores and average cosine similarities based on PANNs embeddings between the set of reference explosion sounds and the synthesized sound set from PronounSE, T-Foley and Stable Audio 2.0. The ± indicates the standard deviation

FAD scores by PANNs Cosine similarity ±SD
Speaker IDTF (Chung et al., 2024)OursSA2 (Evans et al., 2024)TF (Chung et al., 2024)OursSA2 (Evans et al., 2024)
m-0154.0217.8521.350.73±0.070.85±0.080.83±0.06
m-0254.4717.0520.940.74±0.070.86±0.060.84±0.06
m-0348.8517.2820.950.75±0.070.86±0.070.84±0.06
m-0448.5216.6521.460.75±0.080.85±0.070.84±0.06
m-0547.8919.2919.750.75±0.080.86±0.070.85±0.07
f-0156.9419.2822.010.72±0.070.85±0.070.85±0.06
Whole55.9314.0921.150.74±0.070.85±0.070.84±0.06

or Create an Account

Close Modal
Close Modal