An effective speaker adaption using deep learning for the identification of speakers in emergency situation
摘要
Automated speaker identification is an important research topic in recent advanced technologies. This process helps to analyze the speakers in their emergencies. Various existing approaches are used for speaker identification, and the existing systems cannot distinguish background noises such as music and traffic, which can cause signal distortion. Moreover, the existing speaker identification techniques do not have better learning ability and face high computational complexity. Also, they failed to recognize the speaker’s condition because of the inability to extract suitable keywords and noisy speeches. To overcome this issues, the proposed study introduced a new deep-learning mechanism for effective speaker identification in emergency situations. The input audio files are initially collected, and pre-processing is performed to reduce the noise issue using the original speech separation process. The keywords are spotted using the proposed Mel-Frequency Cepstral Coefficients with Binary Weighted Network (MFCC-BWN) from the pre-processed data. This keyword extraction step helps to improve the recognition accuracy because of the significant attributes. Finally, audio files using Adaptive Improved LSTM (AILSTM) identify the speaker’s situation. The simulation result analysis shows that the proposed model outperforms the other comparable methods.