Emotion Recognition in Speech Using Convolutional Neural Networks (CNNs)
摘要
Speech Emotion Recognition (SER) represents a burgeoning technology with the potential to revolutionize human–computer interaction (HCI). In this research paper, we investigate the effectiveness of multiple feature combinations, including Mel frequency cepstral coefficients (MFCC), root mean square error (RMSE), Chroma short-time Fourier transform (STFT), and zero crossing rate (ZCR), to accurately classify emotions into six distinct categories: happy, sad, neutral, fear, anger, and disgust. Our proposed convolutional neural network (CNN)-based approach leverages the fusion of MFCC, RMSE, and ZCR features to achieve remarkable results. We achieved an impressive testing accuracy of 95.52% on evaluating the approach using the CREMA-D dataset. These findings demonstrate the potential of our approach to advance the field of SER and enhance HCI applications with more effective emotion recognition capabilities.