Emotional Speech Corpus: A Review
摘要
Speech Emotion Recognition (SER) has gained significant attention in recent years, with applications in human-machine interaction, healthcare, and customer service. A crucial aspect of SER is the availability of high-quality emotion databases that can train and evaluate emotion recognition models. This paper provides a comprehensive review of thirty speech emotion recognition databases, covering various languages, including English, German, Chinese, Russian, Hindi, and other Indian languages. This study analyze the speech emotion databases based on language, number of speakers, emotions recorded, features extracted, and accuracy achieved. The analysis highlights the dominance of English and Hindi languages in SER research, with a significant number of databases available for these languages. The most recorded emotions are sadness, anger, fear, neutral, happiness, surprise, and disgust. Furthermore, it has been found that most databases extracted Mel-Frequency Cepstral Coefficients (MFCCs) and prosodic features for emotion recognition, with accuracies ranging from 60% to 98%. The review provides valuable insights into the existing SER databases, highlighting the need for more diverse and representative datasets that can improve the performance and generalizability of emotion recognition models.