A Dual ML Framework for Automatic Speech Emotion Recognition
摘要
Automatic Speech Emotion Recognition (SER) assists in improving interaction between human and computer-based systems for emotions detection. Thus, the current study seeks to develop a dual machine learning concept for combination of conventional machine learning with deep learning for accurate SER determination. It concerned on feature extraction and feature selection from the speech signal by Mel-Frequency Cepstrum Coefficients (MFCCs) and Modulation Spectral Features (MSF). These features are given as inputs to the Recurrent Neural Network (RNN) and the Support Vector Machine (SVM) classifier and the performance of both is compared. Besides, Recursive Feature Elimination (RFE) feature selection technique is used to tackle the problem of high dimensionality and potential overfitting, while improving the model’s generalization ability. The proposed approach is verified using the Berlin and Spanish emotional speech databases and the system gets 94% recognition accuracy on the Spanish dataset. As can be observed from the results generated in this study, the integration of traditional and deep learning models that employs efficient feature extraction and selection brings about a vast improvement of the SER system.