A Comparative Study on Speech Emotion Recognition Using Machine Learning
摘要
Emotions can be detected from a person’s speech during communication. The expression of emotions through voice is an ongoing field of research. In this study SAVEE and IEMOCAP datasets were used with regards to the task of speech emotion recognition. There are seven emotions in the SAVEE datasets and four out of eleven emotions in the IEMOCAP dataset which are considered. Features are extracted from the raw audio files, namely ZCR, F0, MFCC, and RMS, and mean of the features are taken. The study shows a comparative analysis on detecting various emotions on both the datasets. The models used are RNN, LSTM, Bi-LSTM, RF, Rotation Forest, and Fuzzy. On the SAVEE dataset, the RF had the maximum accuracy of 76%, followed by Bi-LSTM with 72%. On the IEMOCAP dataset, RF achieves the maximum accuracy of 68 and 67% on the male and female samples, respectively, followed by Bi-LSTM, which achieves 64 and 63% on the male and female samples, respectively. The fuzzy model improved from an accuracy of 38–47%, and Rotation Forest deteriorated from an accuracy of 66–53% on SAVEE to IEMOCAP dataset. A diagnostic user interface is designed to classify the human emotions using the trained models.