错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-featured Speech Emotion Recognition Using Extended Convolutional Neural Network

  • Arun Kumar Dubey,
  • Yogita Arora,
  • Neha Gupta,
  • Sarita Yadav,
  • Achin Jain,
  • Devansh Verma

摘要

There has been a significant increase in recent years in the investigation of emotions expressed via speech signals; this field is known as Speech Emotion Recognition (SER). SER holds immense potential across various applications and serves as a pivotal bridge in enhancing Human-Computer Interaction. However, prevailing challenges such as diminished model accuracy in noisy environments have posed substantial obstacles in this field. To address the scarcity of robust data for SER, we adopted data augmentation techniques, encompassing noise injection, stretching, and pitch modification. Distinguishing our approach from recent literature, we harnessed multiple audio features, including Mel-Frequency Cepstral Coefficients (MFCCs), mel spectrograms, zero crossing rate, root mean square, and chroma. This paper employs Convolutional Neural Networks (CNNs) as the foundation for emotion classification. The Toronto Emotional Speech Set (TESS) and the Ryerson Audio-Visual Data-base of Emotional Speech and Song (RAVDESS) are two well-established datasets that we utilize. The accuracy of our proposed model on the RAVDESS dataset is 72%, and on the TESS dataset, it achieves an impressive 96.62%. These results surpass those of extant models that have been customized for each specific dataset.