Deep-learning-based models have shown significant potential in speech spoof detection, which is crucial to ensuring the authenticity of speech signals. This work aims to expand the knowledge about deep learning-based spoof detection by integrating ResNet50 with linear discriminant analysis (LDA) to reduce the dimensionality. Using the logical access (LA) subset from the ASVspoof 2019 dataset, we generated mel-spectrogram and gammatone spectrogram representations of the speech signals. ResNet50 was used to extract deep features from these spectrograms, and subsequently LDA was applied to reduce feature dimensionality and improve classification accuracy. Our method significantly outperformed the baseline ResNet50 model by reducing the equal error rate (EER) by 43.55% and increasing balanced accuracy by 48.59% for duplicated mel-spectrogram tensor, 8.95% and 15.52% for differentiated mel-spectrogram tensor, and 44.14% and 44.77% for differentiated gammatone spectrogram tensor, respectively. These results demonstrate the effectiveness of combining ResNet50 with gammatone spectrograms and LDA, providing a more robust solution for audio spoof detection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Anti-spoofing Using ResNet50 with Linear Discriminant Analysis for Automatic Speaker Verification

  • Peemapot Uparakool,
  • Waree Kongprawechnon,
  • Noppharut Pipopsophonchai,
  • Aticha Numsonthi,
  • Kavinnat Rawanggij,
  • Napat Assawamanachai,
  • Jessada Karnjana

摘要

Deep-learning-based models have shown significant potential in speech spoof detection, which is crucial to ensuring the authenticity of speech signals. This work aims to expand the knowledge about deep learning-based spoof detection by integrating ResNet50 with linear discriminant analysis (LDA) to reduce the dimensionality. Using the logical access (LA) subset from the ASVspoof 2019 dataset, we generated mel-spectrogram and gammatone spectrogram representations of the speech signals. ResNet50 was used to extract deep features from these spectrograms, and subsequently LDA was applied to reduce feature dimensionality and improve classification accuracy. Our method significantly outperformed the baseline ResNet50 model by reducing the equal error rate (EER) by 43.55% and increasing balanced accuracy by 48.59% for duplicated mel-spectrogram tensor, 8.95% and 15.52% for differentiated mel-spectrogram tensor, and 44.14% and 44.77% for differentiated gammatone spectrogram tensor, respectively. These results demonstrate the effectiveness of combining ResNet50 with gammatone spectrograms and LDA, providing a more robust solution for audio spoof detection.