A hybrid approach to secure automatic speaker verification: integrating clone detection and speaker identification
摘要
The increasing adoption of voice-enabled devices demands robust Automatic Speaker Verification (ASV) systems, especially in sensitive domains like finance and healthcare. However, ASV systems are prone to spoofing attacks, including voice cloning and replay attacks, which compromise their reliability. This paper presents a Secure Automatic Speaker Verification (SASV) system that integrates Linear Frequency Cepstral Coefficients (LFCC) and Inverse Gammatone Cepstral Coefficients (IGTCC) for feature extraction, capturing both spectral and temporal characteristics of audio signals. The system is designed to perform speaker identification and clone detection by differentiating genuine voices from spoofed ones. Additionally, after successful clone detection, the system generates a profile for the detected speaker. For classification, Random Forest and Neural Network models are utilized. The model is evaluated on both Logical Access (LA) and Physical Access (PA) scenarios using the ASVspoof 2019 dataset. This includes attacks such as text-to-speech (TTS), voice conversion (VC), and replay. The results show that the system achieves 96% accuracy in clone detection using a Neural Network and a 4.4% Equal Error Rate (EER) with Random Forest, demonstrating its effectiveness in speaker identification, spoofing detection, and profile generation, thereby enhancing security in ASV applications.