Emotional speaker identification using PCAFCM-deepforest with fuzzy logic
摘要
Voice is perceived as a form of biometrics which communicates valuable and rich information pertinent to an individual, such as his or her identity, gender, accent, age and emotion. Speaker identification denotes the task of identifying speakers based on their intrinsic voice characteristics. This study proposes a text-independent speaker identification system based on principal component analysis (PCA), fuzzy C-means (FCM) along with deepForest called PCAFCM-deepForest. The proposed approach is evaluated under neutral and adverse talking environments. Given this approach, we assessed our proposed model architecture on five benchmark corpora, namely private Arabic Emirati-accented speech dataset, public English dataset; Crowd-sourced emotional multimodal actors dataset (CREMA), public German database; Berlin database of emotional speech (EmoDB), public Chinese and English; emotional speech database (ESD), and public French dataset; public Canadian French emotional (CaFE) speech dataset. Our analysis shows that the performance of speaker identification has been immensely increased (greatly improved) when fuzzy logic and PCA are both applied to the extracted mel-frequency cepstral coefficients (MFCC). Speaker identification performance achieved by the proposed PCAFCM-deepForest is superior to that obtained by deepForest alone, FCM-deepForest as well as convolutional neural network (CNN). Besides, it surpasses the following conventional models: Random forest and support vector machine (SVM). Our findings demonstrate that the attained average speaker identification accuracy is equivalent to 98.20% using the Emirati database; an average performance which outperforms the existing frameworks. Moreover, PCAFCM-deepForest is fine-tuned using the grid search algorithm, and the achieved complexity is much less than that of CNN.