Table 6.

Detection of AI-generated voices

NameFeatures/inputsModel/approachAttack types coveredEvaluation data set(s)Key metricsYear
CQCC-GMM (Todisco et al., 2019)CQCC (handcrafted)Gaussian mixture modelVC, statistical TTSASVspoof 2015/2019EER, t-DCF2019
LFCC-GMM (Todisco et al., 2019)LFCC (handcrafted)Gaussian mixture modelVC, statistical TTSASVspoof 2019EER, t-DCF2019
LFCC-LCNN (Lavrentyeva et al., 2019)LFCC + spectrogramLight CNNVC, TTS (logical access)ASVspoof 2019/2021EER, t-DCF2019
RawNet2 (Tak et al., 2021a)Raw waveformCNN-GRU end-to-endTTS, VC, replayASVspoof 2019/2021EER2021
AASIST (Jung et al., 2021)Spectrogram + graph featuresSpectro-temporal graph attention networkVC, TTS, replay (logical/physical access)ASVspoof 2021EER, t-DCF2021
RawNet3 (Tak et al., 2021a)Raw waveformEnhanced CNN-GRUTTS, VC, replayASVspoof 2021EER2021
WavLM (Chen et al., 2022a)Self-supervised embeddingsTransformer-based SSL modelTTS, VC, unseen neural synthesisLibriSeVoc, WaveFake, ASVspoof 2021EER, accuracy2022
XLS-R (Babu et al., 2021)Cross-lingual embeddingsSelf-supervised multilingual transformerTTS, VC (cross-lingual)WaveFake, ASVspoof 2021EER, accuracy2021
RawNet2 Vocoder (Sun et al., 2023)Raw waveform + vocoder IDMultitask RawNet2Neural vocoder tracesLibriSeVoc, WaveFake, ASVspoof 2019EER, robustness tests2023

or Create an Account

Close subscription notice
Close access options