Speech Emotion Recognition (SER) has gained significant attention in recent years, with applications in human-machine interaction, healthcare, and customer service. A crucial aspect of SER is the availability of high-quality emotion databases that can train and evaluate emotion recognition models. This paper provides a comprehensive review of thirty speech emotion recognition databases, covering various languages, including English, German, Chinese, Russian, Hindi, and other Indian languages. This study analyze the speech emotion databases based on language, number of speakers, emotions recorded, features extracted, and accuracy achieved. The analysis highlights the dominance of English and Hindi languages in SER research, with a significant number of databases available for these languages. The most recorded emotions are sadness, anger, fear, neutral, happiness, surprise, and disgust. Furthermore, it has been found that most databases extracted Mel-Frequency Cepstral Coefficients (MFCCs) and prosodic features for emotion recognition, with accuracies ranging from 60% to 98%. The review provides valuable insights into the existing SER databases, highlighting the need for more diverse and representative datasets that can improve the performance and generalizability of emotion recognition models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotional Speech Corpus: A Review

  • Krishna Rohilla,
  • Aaditya Sharma,
  • Nilakshi,
  • Nivedita Palia

摘要

Speech Emotion Recognition (SER) has gained significant attention in recent years, with applications in human-machine interaction, healthcare, and customer service. A crucial aspect of SER is the availability of high-quality emotion databases that can train and evaluate emotion recognition models. This paper provides a comprehensive review of thirty speech emotion recognition databases, covering various languages, including English, German, Chinese, Russian, Hindi, and other Indian languages. This study analyze the speech emotion databases based on language, number of speakers, emotions recorded, features extracted, and accuracy achieved. The analysis highlights the dominance of English and Hindi languages in SER research, with a significant number of databases available for these languages. The most recorded emotions are sadness, anger, fear, neutral, happiness, surprise, and disgust. Furthermore, it has been found that most databases extracted Mel-Frequency Cepstral Coefficients (MFCCs) and prosodic features for emotion recognition, with accuracies ranging from 60% to 98%. The review provides valuable insights into the existing SER databases, highlighting the need for more diverse and representative datasets that can improve the performance and generalizability of emotion recognition models.