Speech Emotion Recognition Using Deep Learning Algorithm on RAVDESS Dataset
摘要
The processing of speech signals is an active topic of research, the most important medium for the communication of information between human beings, and the most effective medium for human–computer interaction (HCI). Identification of human actions’ emotions based on a voice signal, often known as speech emotion recognition (SER), is an emergent domain of research in the field of HCI that has numerous real-time claims. The learning features which may contain information that is relevant and discriminative such as high-level deep features are essential to the performance of an effective SER system. In this research, we suggested the multilayer perceptron (MLP) classifier as the method for making the ultimate prediction. In the course of this investigation, the Mel Frequency Cepstral Coefficient (MFCC), Chroma, and Mel properties are all taken into consideration. In this particular research endeavor, the MFCC, Chroma, and Mel spectrogram were the primary techniques that were utilized for the purpose of feature extraction. These features are extracted with the help of Librosa's built-in functions, and after that, they are placed in arrays so that they may be categorized and analyzed. We trained our proposed system with the RAVDESS emotional speech corpora that serve as a benchmark, and when we assessed the performance of the prediction, we were able to get recognition rates of 87.9%. The efficiency and significance of the suggested system can be seen clearly through the performance of the system.