Comprehensive Analysis of Probabilistic and Machine Learning Models for Speech Emotion Recognition
摘要
The recognition of sentiments in language is crucial for enhancing human–machine collaboration and affective computing, with its applications ranging from virtual assistants to mental health diagnostics. This research investigates the effectiveness of a spectrum of machine learning algorithms in SER applications, aiming to push the boundaries of emotion detection accuracy and applicability. Leveraging a diverse collection of datasets counting the Berlin Speech Emotive Database, Collaborating Emotive Dyadic Motion Capture (IEMOCAP), and the RAVDESS dataset formed the basis of the research, ensuring a comprehensive exploration of emotional expressions across different contexts and demographics. A wide range of machine learning algorithms was employed, from statistical analytical techniques like Gaussian Mixture Models (GMMs) and Hidden Markov Models (HMMs) to cutting-edge neural network architectures. Support Vector Machines (SVMs) were utilized for delineating crucial boundaries, while the temporal capabilities of Recurrent Neural Networks (RNNs) and the enduring memories embedded in Long Short-Term Memory networks (LSTMs) were explored. Additionally, Convolutional Neural Networks (CNNs) brought their pixelated expertise in analyzing spectrograms and other visual representations of speech. Ensemble methods such as Light Gradient Boosting Machines, Random Forest Classifiers, Extra Trees Classifiers, Gradient Boosting Classifiers, and Multi-Layer Perceptron Classifiers were also employed, weaving tapestries of predictive power and robustness. Through meticulous evaluation metrics incorporating metrics such as accuracy, precision, recall, and F1-score, this study offers a thorough examination of the strengths and weaknesses of each algorithm. Notably, results indicate the prominence of sequential models, particularly LSTM, in capturing nuanced emotional patterns, attaining a level of accuracy of 93.60% on the (IEMOCAP) dataset. Novel insights include the effective integration of diverse datasets and the comparison of various machine learning algorithms, shedding light on their performance in real-world SER contexts. This research contributes to the advancement of the SER field, offering valuable guidance for algorithm selection and optimization in practical applications, thereby opening avenues for human–computer interactions that are more empathetic and intuitive.