Automatic speaker identification system based on MLP network and deep learning in the presence of severe interference
摘要
Automatic Speaker Identification (ASI) is so crucial for security. Current ASI systems perform well in quiet and clean surroundings. However, in noisy situations, the robustness of an ASI system against additive noise and interference is a crucial factor. An investigation of the impact of interference on ASI system performance is presented in this paper, which introduces algorithms for achieving high ASI system performance. The objective is to resist the interference of various forms. This paper presents two models for the ASI task in the presence of interference. The first one depends on Normalized Pitch Frequency (NPF) and Mel-Frequency Cepstral Coefficients (MFCCs) as extracted features and Multi-Layer Perceptron (MLP) as a classifier. In this model, we investigate the utilization of a Discrete Transform (DT), such as Discrete Wavelet Transform (DWT), Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST), to increase the robustness of extracted features against different types of degradation through exploiting the sub-band decomposition characteristics of DWT and the energy compaction property of DCT and DST. This is achieved by extracting features directly from contaminated speech signals in addition to features extracted from discrete transformed signals to create hybrid feature vectors. The enhancement techniques, such as Spectral Subtraction (SS), Winer Filter, and adaptive Wiener filter, are used in a preprocessing stage to eliminate the effect of the interference on the ASI system. In the second model, we investigate the utilization of Deep Learning (DL) based on a Convolutional Neural Network (CNN) with speech signal spectrograms and their Radon transforms to increase the robustness of the ASI system against interference effects. One of this paper goals is to introduce a comparison between the two models and build a more robust ASI system against severe interference. The experimental results indicate that the two proposed models lead to satisfactory results, and the model based on CNN consumes less time than that of the model based on MLP, which requires much training epochs.