NRASV: Noise Robust ASV System for Audio Replay Attack Detection
摘要
Recent advancements in researching countermeasure systems against audio replay attacks have notably bolstered the resilience of Automatic Speaker Verification (ASV) Systems. However, current countermeasures face challenges in adapting to unknown attack variations. This limitation stems from the failure to capture context-specific details from the speech data, a crucial factor leading to the suboptimal performance of these systems against novel threats. To address this, our proposed anti-replay system integrates Gammatone features with t-distributed Stochastic Neighbor Embedding (t-SNE) for dimensionality reduction. In the classification phase, we employ diverse machine learning algorithms including Support Vector Machine (SVM), Random Forest (RF), K Nearest Neighbor (KNN), and Gradient Boost (Gboost). Our evaluation, conducted on the VSDC dataset in both controlled and noisy environments, introduces babble noise at varying levels (15dB, 10dB, 5dB, and 0dB). Remarkably, our model, GTCC-t-SNE with SVM, attains exceptional performance with Equal Error Rates (EER) of 2.2%, 3.3%, 3.8%, and 4.4% for the respective noise levels, demonstrating its robustness against audio replay attacks.