Combined Bidirectional Long Short-Term Memory and Mel-Frequency Cepstral Coefficients with Convolution Neural Network Using Triplet Loss for Speaker Recognition
摘要
In recent years, there have been significant breakthroughs in neural network technology within the field of speaker recognition, encompassing various applications such as word classification, emotion recognition, and speaker identification. This paper introduces an innovative method for speaker recognition aimed at enhancing accuracy. The method, termed bidirectional long short-term memory combined with mel-frequency cepstral coefficients and convolution neural network features using triplet loss for speaker recognition (BLSTM-MFCCCNN-TL), employs dual-factor features as the input layer of the model. Compared to traditional GMM-HMM and BLSTM-MFCC models, the new method not only provides more accurate recognition capabilities but also exhibits faster computational speeds. Consequently, these advancements significantly enhance the overall performance of speaker recognition while improving computational efficiency.