<p>Automated speaker recognition has recently emerged as a hot and demanding area of study. In this work, we provide a new approach to speaker detection and verification that makes use of attention mechanisms in conjunction with deep multidimensional acoustic feature acquisition. This approach is based on several deep convolutional neural networks (CNN). In order to get deep, segment-level speaker-specific characteristics from the source speech signal, the suggested method first creates three separate audio components for multi-CNNs: one for training with an unprocessed waveform in one dimension, one for training with a time–frequency Mel-spectrogram in two dimensions, and one for training with dynamic time–space features. The utterance-level analysis outcomes are generated by applying global average pooling to the segment-level outputs acquired from 1D, 2D, and 3D CNN models. Finally, for speaker identification, an attention-based approach skillfully combines features from the three streams to incorporate various outcomes from utterance-level categorization. The extracted deep multimodal speaker properties are demonstrated to be mutually beneficial, allowing for their integration in an attention-based fusion network to yield substantially enhanced performance. In order to test the suggested scheme, we used a number of conventional and real-time audio datasets. The suggested attention-based multi-dimensional fused-feature convolutional neural network (AMDF-CNN) reduces the speaker misclassification error rate by 2.52% when tested against baseline approaches, according to the experimental results. With an impressive identification rate of 97.59%, the AMDF-CNN speaker identification model performed well in the experiments. All the while, we put the model through its paces under different kinds of noise to see how reliable it is. Experimental results show that the suggested strategy outperforms state-of-the-art schemes with a reliability of more than 85%. Relevance of the work: The integration of artificial intelligence (AI) and deep learning (DL) systems has great promise for the future of investigative disciplines, automated processes, and verification in the field of speaker detection. Improving the detection and verification procedure with an AMDF-CNN can help with disregarding various difficulties like ambient noise, false positives, and more. It is possible to build a disease-treating mechanism called raga detection and treatment by expanding this technique.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-based multi dimension fused-feature convolutional neural network framework for speaker recognition

  • V. Karthikeyan,
  • S. Suja Priyadharsini,
  • K. Balamurugan

摘要

Automated speaker recognition has recently emerged as a hot and demanding area of study. In this work, we provide a new approach to speaker detection and verification that makes use of attention mechanisms in conjunction with deep multidimensional acoustic feature acquisition. This approach is based on several deep convolutional neural networks (CNN). In order to get deep, segment-level speaker-specific characteristics from the source speech signal, the suggested method first creates three separate audio components for multi-CNNs: one for training with an unprocessed waveform in one dimension, one for training with a time–frequency Mel-spectrogram in two dimensions, and one for training with dynamic time–space features. The utterance-level analysis outcomes are generated by applying global average pooling to the segment-level outputs acquired from 1D, 2D, and 3D CNN models. Finally, for speaker identification, an attention-based approach skillfully combines features from the three streams to incorporate various outcomes from utterance-level categorization. The extracted deep multimodal speaker properties are demonstrated to be mutually beneficial, allowing for their integration in an attention-based fusion network to yield substantially enhanced performance. In order to test the suggested scheme, we used a number of conventional and real-time audio datasets. The suggested attention-based multi-dimensional fused-feature convolutional neural network (AMDF-CNN) reduces the speaker misclassification error rate by 2.52% when tested against baseline approaches, according to the experimental results. With an impressive identification rate of 97.59%, the AMDF-CNN speaker identification model performed well in the experiments. All the while, we put the model through its paces under different kinds of noise to see how reliable it is. Experimental results show that the suggested strategy outperforms state-of-the-art schemes with a reliability of more than 85%. Relevance of the work: The integration of artificial intelligence (AI) and deep learning (DL) systems has great promise for the future of investigative disciplines, automated processes, and verification in the field of speaker detection. Improving the detection and verification procedure with an AMDF-CNN can help with disregarding various difficulties like ambient noise, false positives, and more. It is possible to build a disease-treating mechanism called raga detection and treatment by expanding this technique.