Siam-CNNNet: A Novel Fusion of Siamese Network and Convolutional Neural Networks Based on Mel-Frequency Cepstral Coefficients for Audio Deepfake Detection
摘要
Recent advancements in Artificial Intelligence, particularly in Generative Artificial Intelligence (GAI), have led to the emergence of increasingly realistic audio deep-fake technology. This technology enables manipulating audio content associated with a source's audio. The recent development of audio deepfakes presents significant concerns as they introduce the creation of highly authentic spoken words, posing a threat through potential misuse for impersonation or dissemination of misinformation. Effective detection methods for identifying such manipulations must exhibit robustness against attacks employing techniques not explicitly accounted for during training, necessitating characteristics such as strong generalization and stability. This paper proposes a novel Audio Deepfake Detection (ADD) model based on Convolutional Neural Networks (CNN) for extracting audio features from Mel-Frequency Cepstral Coefficients (MFCCs). Then, a Siamese network is employed to recognize real and fake audio samples. Experimental results conducted on two benchmark datasets, ASV Spoof 2019 and DEEP-VOICE, demonstrate the efficiency of the proposed model. The results demonstrate high accuracy, achieving 99.9% and 96.9% accuracy rates for ASV Spoof 2019 and DEEP-VOICE, respectively.