Deep normalization for light SpineNet speaker anti-spoofing systems
摘要
Despite their impressive performance in controlled conditions, current speaker recognition systems still face challenges related to the diversity of real-world situations, including unpredictable noisy conditions and spoofing attacks. This paper presents a novel approach that optimizes the deployment of automatic speaker verification spoofing countermeasures. An innovative normalization process is proposed to adapt Light SpineNet-based countermeasure vectors for this optimization in conjunction with the probabilistic linear discriminant analysis (PLDA) scoring method. Three normalization techniques –maximum Gaussianality discriminative normalization flow (MG-DNF), maximum likelihood discriminative normalization flow (ML-DNF), and variational autoencoder regularization (VAE)– are assessed by using the logical access evaluation dataset of the ASVspoof 2021 challenge edition. This dataset includes diverse transmission artifacts and realistic conditions, enabling the evaluation of the ability of the normalized Light SpineNet-based countermeasures embedding to prevent spoofing attacks. The results showed the effectiveness of the introduced normalization approach within the LSpineNet-based anti-spoofing system. The LSpineNet49-GM-DNF countermeasure embedding achieved the best performance compared to DNF-, VAE-based, and current state-of-the-art systems.