Effect of Interference on Text-independent Speaker Recognition Based on Deep Learning
摘要
In the evolving landscape of Speaker Recognition (SR), the majority of research traditionally focuses on scenarios devoid of ambient noise or interference, largely adhering to text-dependent protocols, where the identity is confirmed through specific verbal cues. This approach, while methodologically efficient under controlled conditions, starkly contrasts with the complex and unpredictable nature of real-world environments in the presence of various acoustic disturbances. Addressing this gap, our study is concerned with the exploration of text-independent SR under realistic interference conditions, aiming to authenticate speakers based on the intrinsic characteristics of their speech, independent of the linguistic content conveyed. Central to our methodology is the strategic employment of advanced signal separation, serving as a critical pre-processing step to reduce the impact of external noise and interference. This work facilitates a more accurate and efficient identification process, pivotal for text-independent SR systems. Our framework harnesses the computational prowess of two leading-edge Deep Learning (DL) architectures: the Long-Short Term Memory Recurrent Neural Network (LSTM-RNN) and the Deep Convolutional Neural Network (CNN). Through a meticulous comparative analysis, we evaluate these models' efficacy in various interference landscapes, benchmarking their performance against established conventional systems. A novel aspect of our research is the application of signal separation in different domains, yielding significant enhancements in SR accuracy, notably within the time domain. Empirical results of experiments reveal that a three-layer CNN configuration achieves a high recognition rate of 97.33%, with the LSTM-RNN model closely following with 96%. These models not only surpass the benchmarks set by traditional two-stage attention models, including Time Delay Neural Networks (TDNNs) and CNNs, achieving 91.1% and 92%, respectively, but also highlight the potential of our approach. Furthermore, our study delves into the comparative efficiency of the two DL models under scrutiny, elucidating their respective strengths and limitations in the context of SR. The profound impact of domain-specific signal separation as a pre-processing measure emerges as a pivotal finding, markedly enhancing the system resilience and accuracy in determining speaker identities amidst interference. In conclusion, our research ensures the capability of the proposed text-independent SR system, augmented by signal separation and DL methodologies, to navigate and triumph over the challenges posed by interference in real-world settings. This study not only sets a new benchmark in the field of SR, but also opens directions for future research, particularly in enhancing the robustness and adaptability of SR systems to operate in different environmental conditions.