Novel WaveELU for Spoof Classification in ASV System Using Fusion Features
摘要
Biometric authentication plays a crucial role in protecting data on devices. Automatic Speaker Verification (ASV), a biometric authentication system that grants access through voice, has garnered significant attention in recent years. However, ASV systems are vulnerable to various attacks, including voice conversion (VC), speech synthesis (SS), replay, mimicry, and Twin attacks. The recent works of the ASV system have not attempted mimicry and twin attacks. In addition, the systems cannot classify each attack if the attack’s occurrence is uncertain. This manuscript proposes a novel deep-learning WaveELU model to classify spoof attacks as a multiclass classification model, employing a fused feature extraction technique for better accuracy. Feature extraction technique includes fusing relevant features from Mel-Frequency Cepstral Coefficients (MFCC) and spectrogram features. The system achieves better accuracy by using these fused features for the proposed model’s training and testing phases. It is observed from the simulation results that the novel WaveELU model provides an Equal Error Rate (EER) of 8.62% and 2.63% for spectrogram and MFCC features, respectively. The proposed model provides a reduced EER of 1.37% through the fusion of features. Furthermore, the system’s accuracy improves by a factor of 2.21% compared to the existing model.