Cross-Linguistic Speech Emotion Detection Using Convolutional Neural Networks
摘要
Speech Emotion Recognition (SER) is a complex and innovative computer method created to identify and group human audio signals into emotions. This approach enables machines to interact with humans’ feelings. This paper explains a practically feasible SER, which is the analysis of speech to determine emotional content through a neural network algorithm. We initially used around 7,500 the audio samples as our dataset, and then the dataset was further magnified using audio augmentation process to 20,000. Our system primarily employs a Convolutional Neural Network using Mel-Frequency Cepstral Coefficients features. We evaluated the performance of this model on datasets for both Azerbaijani and English. Our findings indicate that the language of the dataset significantly influences the effectiveness of emotion detection.