Advanced Speech Emotion Recognition in Malayalam Accented Speech: Analyzing Unsupervised and Supervised Approaches
摘要
Emotion Recognition for accented Malayalam speech poses a significant challenge due to the complexities inherent in the language. In this study, we explore the effectiveness of utilizing an unsupervised approach by incorporating various clustering algorithms and supervised approach by incorporating convolutional neural networks to this specific dataset. The performance of the various approaches has been evaluated using Silhouette Score for the unsupervised approach and accuracy, precision, recall and F1-score for the supervised approaches. Affinity Propagation, Optics Clustering, Mean Shift Clustering, Agglomerative Clustering, Gaussian Mixture Model Clustering, Balanced Iterative Reducing and Clustering using Hierarchies, Consensus Clustering and Ensembled Clustering techniques were adopted employing unsupervised clustering techniques. The Affinity Propagation method performed well with a Silhouette Score of 0.5255 which characterizes a superior cluster quality resulting in well-defined and distinct clusters. Mean Shift Clustering and OPTICS both performed well, with scores of 0.2511 and 0.4029, respectively. This implies that discrete and well-defined clusters have been formed. Furthermore, experiments were conducted using Ensemble Clustering (Majority Voting), which achieved a moderate degree of cluster distinctiveness with a score of 0.2399. These results provide an insightful viewpoint on the possible benefits of ensemble approaches in this situation. The outcomes of our experiment demonstrate how well different clustering techniques work when it comes to classifying emotions in Malayalam speech with accents. The speech dataset used in the study is augmented using noise addition, shifting, and stretching. The supervised approach of emotion classification using CNN yielded a high accuracy of 99.9%. Additionally, the experiment's recall, accuracy, and F1-score were all high.