Detection of Audio Spoofing Attack Using Integrated Hand-Crafted Features with LSTM
摘要
Spoofing attacks, advanced text-to-speech (TTS) mechanisms, and voice conversion (VC) technologies pose significant security risks to Automatic Speaker Verification (ASV) systems by generating realistic-sounding synthetic speech and spoofed voices. To protect ASV systems from unauthorized access, it is essential to detect such synthetic speech and spoofed voices effectively. An Audio Spoof Detection (ASD) system typically comprises two main components: front-end feature extraction and back-end classification. In this paper, we introduce a novel approach to address this challenge by combining hand-crafted features with graph-based features during the front-end feature extraction stage. Specifically, we utilize Acoustic Ternary Pattern (ATP), Mel Frequency Cepstral Coefficient (MFCC), Gammatone Cepstral Coefficient (GTCC), and a new graph-based feature called Graph Frequency Cepstral Coefficient (GFCC). For the back-end classification, we employ Long Short-Term Memory (LSTM) networks. To evaluate the effectiveness of our proposed system, we develop six distinct configurations. The first three systems use ATP, MFCC, and GTCC features individually, while the remaining three systems combine GFCC with ATP, MFCC, and GTCC features, respectively. We conduct our experiments using the ASVspoof 2019 Physical Access (PA) evaluation datasets.