<p>Voiceprint recognition technology, as a form of biometric identification, holds extensive potential applications in the realm of security authentication. However, practical implementations often encounter challenges in cross-scenario and cross-channel recognition, resulting in decreased recognition accuracy. To address this issue, a model based on deep learning and dual-channel voiceprint recognition is proposed in this paper. Firstly, a Distance-weighted Sub-center Arcface Loss (DWLoss) utilizing pole-zero distance weighting is designed to train the voiceprint recognition model, aimed at enhancing the model’s generalization capability and robustness. Secondly, an Equal Channel Attention (ECA) feature channel weighting module for deep feature extraction without dimension reduction is employed, combined with Probabilistic Linear Discriminant Analysis (PLDA) for channel compensation and speaker scoring to tackle the cross-channel recognition problem. Finally, based on the aforementioned designs, the SE-Res2Net-DWLoss-TDNN and ECA-Res2Net-TDNN-PLDA dual-channel voiceprint recognition models are presented. The weight matrix is trained to combine the scoring scores of the two channels, thereby obtaining a comprehensive speaker similarity score. Experimental validation demonstrates the Synergistic nature of the two channels during the recognition process, as well as the superiority of the model in improving recognition performance and robustness. Experimental results indicate that compared to traditional models, this model exhibits significant advantages in noise resistance and recognition accuracy. This study contributes beneficial advancements and enhancements to cross-scenario and cross-channel recognition within the realm of voiceprint recognition technology.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mitigating cross-scenario challenges in voiceprint recognition: a dual-channel approach

  • Huajun Zhang,
  • Shuqi Wang

摘要

Voiceprint recognition technology, as a form of biometric identification, holds extensive potential applications in the realm of security authentication. However, practical implementations often encounter challenges in cross-scenario and cross-channel recognition, resulting in decreased recognition accuracy. To address this issue, a model based on deep learning and dual-channel voiceprint recognition is proposed in this paper. Firstly, a Distance-weighted Sub-center Arcface Loss (DWLoss) utilizing pole-zero distance weighting is designed to train the voiceprint recognition model, aimed at enhancing the model’s generalization capability and robustness. Secondly, an Equal Channel Attention (ECA) feature channel weighting module for deep feature extraction without dimension reduction is employed, combined with Probabilistic Linear Discriminant Analysis (PLDA) for channel compensation and speaker scoring to tackle the cross-channel recognition problem. Finally, based on the aforementioned designs, the SE-Res2Net-DWLoss-TDNN and ECA-Res2Net-TDNN-PLDA dual-channel voiceprint recognition models are presented. The weight matrix is trained to combine the scoring scores of the two channels, thereby obtaining a comprehensive speaker similarity score. Experimental validation demonstrates the Synergistic nature of the two channels during the recognition process, as well as the superiority of the model in improving recognition performance and robustness. Experimental results indicate that compared to traditional models, this model exhibits significant advantages in noise resistance and recognition accuracy. This study contributes beneficial advancements and enhancements to cross-scenario and cross-channel recognition within the realm of voiceprint recognition technology.