Abstract <p>Speech is the most effective way for people to communicate with each other. It may transmit rich and important information including accent, gender, emotion, and distinct identity. A key property of a good automatic speaker recognition (ASR) system is being able to model the uncertainty and variation of the utterances of the same speaker. Speaker recognition under noisy and unconstrained conditions is an extremely challenging. Machine learning is one of the most widely used systems for speaker identification. But achieving better accuracy is difficult due to the inaccurate extraction of features. To overcome these drawbacks, optimized Hopfield neural network (HNN) is developed. Raw audio signals are preprocessed using Notch filter and adaptive thresholding to enhance the quality of the original signal. Then, extracting the features from the preprocessed signal using bark frequency cepstral coefficients (BFCC) as well as generalized perceptual linear prediction (GPLP). In the process of feature extraction, preprocessed signal is decomposed into frequency subbands and further normalized to obtain features coefficients. Extracted features are given as input for the classification process. Here, Hopfield neural network is employed for identifying the speakers in noisy environment. Lotus effect optimization algorithm (LEOA) algorithm is utilized for tuning the hyper parameters in HNN which is used to enhance the classifier performance. Proposed model achieved 94<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11974_2025_8419_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <!--OptelIns2570017Nayak-m1--> </InlineEquation> accuracy, mean absolute percentage error 4.31<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11974_2025_8419_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <!--OptelIns2570017Nayak-m2--> </InlineEquation> and recall 94<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11974_2025_8419_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <!--OptelIns2570017Nayak-m3--> </InlineEquation>. Thus, the optimized Hopfield neural network is the best option for identifying the speakers in the noisy environment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hopfield Neural Network with Lotus Effect Optimization for Automatic Speaker Recognition in Noisy Environment

  • Sabita Nayak,
  • Imteyaz Ahmad,
  • Pawan Kumar

摘要

Abstract

Speech is the most effective way for people to communicate with each other. It may transmit rich and important information including accent, gender, emotion, and distinct identity. A key property of a good automatic speaker recognition (ASR) system is being able to model the uncertainty and variation of the utterances of the same speaker. Speaker recognition under noisy and unconstrained conditions is an extremely challenging. Machine learning is one of the most widely used systems for speaker identification. But achieving better accuracy is difficult due to the inaccurate extraction of features. To overcome these drawbacks, optimized Hopfield neural network (HNN) is developed. Raw audio signals are preprocessed using Notch filter and adaptive thresholding to enhance the quality of the original signal. Then, extracting the features from the preprocessed signal using bark frequency cepstral coefficients (BFCC) as well as generalized perceptual linear prediction (GPLP). In the process of feature extraction, preprocessed signal is decomposed into frequency subbands and further normalized to obtain features coefficients. Extracted features are given as input for the classification process. Here, Hopfield neural network is employed for identifying the speakers in noisy environment. Lotus effect optimization algorithm (LEOA) algorithm is utilized for tuning the hyper parameters in HNN which is used to enhance the classifier performance. Proposed model achieved 94 \(\%\) accuracy, mean absolute percentage error 4.31 \(\%\) and recall 94 \(\%\) . Thus, the optimized Hopfield neural network is the best option for identifying the speakers in the noisy environment.