Assessing random forest performance in low resource speech emotion recognition
摘要
In human-computer interaction (HCI), speech emotion recognition (SER) is a pivotal technology that enables machines to decipher human emotions. This study delves into the efficacy of employing a random forest (RF) classifier for SER within Urdu speech, a notably low-resource language. By focusing on the three primary emotions—happiness, sadness, and anger, and leveraging Mel-frequency cepstral coefficients (MFCCs) for feature extraction, we have meticulously crafted a model with an impressive validation accuracy of 94.53%. This endeavor highlights the robustness of the RF classifier and represents a significant step towards empathetic artificial intelligence, particularly in improving digital user experiences through emotional understanding. Moreover, our examination highlights the key features of MFCCs and concentrating on their crucial role in emotion discrepancy. This research endeavors the viability and realism of RF classifiers in paving the way for emotionally subtle artificial intelligence (AI) systems, even under the constrictions of resource-deprived languages for instance Urdu. In the future, we hope to broaden the scope of our research to include a wider spectrum of emotions and investigate how other datasets affect our model’s functionality, creating new avenues for the advancement of SER technology.