Detection of AI-generated voices
| Name | Features/inputs | Model/approach | Attack types covered | Evaluation data set(s) | Key metrics | Year |
|---|---|---|---|---|---|---|
| CQCC-GMM (Todisco et al., 2019) | CQCC (handcrafted) | Gaussian mixture model | VC, statistical TTS | ASVspoof 2015/2019 | EER, t-DCF | 2019 |
| LFCC-GMM (Todisco et al., 2019) | LFCC (handcrafted) | Gaussian mixture model | VC, statistical TTS | ASVspoof 2019 | EER, t-DCF | 2019 |
| LFCC-LCNN (Lavrentyeva et al., 2019) | LFCC + spectrogram | Light CNN | VC, TTS (logical access) | ASVspoof 2019/2021 | EER, t-DCF | 2019 |
| RawNet2 (Tak et al., 2021a) | Raw waveform | CNN-GRU end-to-end | TTS, VC, replay | ASVspoof 2019/2021 | EER | 2021 |
| AASIST (Jung et al., 2021) | Spectrogram + graph features | Spectro-temporal graph attention network | VC, TTS, replay (logical/physical access) | ASVspoof 2021 | EER, t-DCF | 2021 |
| RawNet3 (Tak et al., 2021a) | Raw waveform | Enhanced CNN-GRU | TTS, VC, replay | ASVspoof 2021 | EER | 2021 |
| WavLM (Chen et al., 2022a) | Self-supervised embeddings | Transformer-based SSL model | TTS, VC, unseen neural synthesis | LibriSeVoc, WaveFake, ASVspoof 2021 | EER, accuracy | 2022 |
| XLS-R (Babu et al., 2021) | Cross-lingual embeddings | Self-supervised multilingual transformer | TTS, VC (cross-lingual) | WaveFake, ASVspoof 2021 | EER, accuracy | 2021 |
| RawNet2 Vocoder (Sun et al., 2023) | Raw waveform + vocoder ID | Multitask RawNet2 | Neural vocoder traces | LibriSeVoc, WaveFake, ASVspoof 2019 | EER, robustness tests | 2023 |
| Name | Features/inputs | Model/approach | Attack types covered | Evaluation data set(s) | Key metrics | Year |
|---|---|---|---|---|---|---|
| CQCC-GMM ( | Gaussian mixture model | VC, statistical | ASVspoof 2015/2019 | EER, t-DCF | 2019 | |
| LFCC-GMM ( | Gaussian mixture model | VC, statistical | ASVspoof 2019 | EER, t-DCF | 2019 | |
| LFCC-LCNN ( | Light | VC, | ASVspoof 2019/2021 | EER, t-DCF | 2019 | |
| RawNet2 ( | Raw waveform | CNN-GRU end-to-end | TTS, VC, replay | ASVspoof 2019/2021 | 2021 | |
| Spectrogram + graph features | Spectro-temporal graph attention network | VC, TTS, replay (logical/physical access) | ASVspoof 2021 | EER, t-DCF | 2021 | |
| RawNet3 ( | Raw waveform | Enhanced CNN-GRU | TTS, VC, replay | ASVspoof 2021 | 2021 | |
| WavLM ( | Self-supervised embeddings | Transformer-based | TTS, VC, unseen neural synthesis | LibriSeVoc, WaveFake, ASVspoof 2021 | EER, accuracy | 2022 |
| XLS-R ( | Cross-lingual embeddings | Self-supervised multilingual transformer | TTS, | WaveFake, ASVspoof 2021 | EER, accuracy | 2021 |
| RawNet2 Vocoder ( | Raw waveform + vocoder | Multitask RawNet2 | Neural vocoder traces | LibriSeVoc, WaveFake, ASVspoof 2019 | EER, robustness tests | 2023 |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.