Designing a universal model for human emotion recognition from speech is challenging due to the variability, subjectivity, and dynamic nature of human emotion expressions. Traditional 2D CNN models rely on singular feature extraction techniques using mel-spectrograms as a 2D matrix (e.g., 64 \(\times \) 128) while 3D CNN offers potential by capturing spatial and temporal features from speech. Aiming to enhance accuracy and robustness in human emotion recognition from speech, we have proposed a 3D CNN model with multi-feature fusion by incorporating three distinct feature extraction techniques such as Mel-frequency Cepstral Coefficients (MFCC), Chroma Shift, and a Mel-spectrogram. The proposed model captures a more comprehensive representation of the underlying emotion dynamics from speech. Instead of flattening the features into a 2D matrix, the proposed model extends the architecture into the 3D, stacking the features along with the z-axis (e.g., 64 \(\times \) 128 \(\times \) 3) fostering a more nuanced understanding of emotional cues in speech. The model is trained and tested using the substantial dataset of 7,000 audio clips from the SUBESCO dataset containing 7 emotion classes, 1,440 audio clips from the RAVDESS dataset containing 8 emotion classes, and 8,440 audio clips from the combined dataset of SUBESCO and RAVDESS datasets containing 8 emotion classes. The proposed model achieves recognition accuracy of 96.61% for the SUBESCO dataset, 89.24% for the RAVDESS dataset, and 94.16% for the combined dataset maintaining a training and testing dataset ratio of 80:20 for each case. The performance results of the proposed model show the significant improvements of the state-of-the-art models in achieving accurate and reliable emotion recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A 3D CNN Model with Multi-feature Fusion for Enhancing Human Emotion Recognition from Speech

  • Hamad Ismail,
  • Sk. Nahid,
  • Md. Zahid Hasan,
  • Md. Parvez Hossain,
  • Muhammad Aminur Rahaman

摘要

Designing a universal model for human emotion recognition from speech is challenging due to the variability, subjectivity, and dynamic nature of human emotion expressions. Traditional 2D CNN models rely on singular feature extraction techniques using mel-spectrograms as a 2D matrix (e.g., 64 \(\times \) 128) while 3D CNN offers potential by capturing spatial and temporal features from speech. Aiming to enhance accuracy and robustness in human emotion recognition from speech, we have proposed a 3D CNN model with multi-feature fusion by incorporating three distinct feature extraction techniques such as Mel-frequency Cepstral Coefficients (MFCC), Chroma Shift, and a Mel-spectrogram. The proposed model captures a more comprehensive representation of the underlying emotion dynamics from speech. Instead of flattening the features into a 2D matrix, the proposed model extends the architecture into the 3D, stacking the features along with the z-axis (e.g., 64 \(\times \) 128 \(\times \) 3) fostering a more nuanced understanding of emotional cues in speech. The model is trained and tested using the substantial dataset of 7,000 audio clips from the SUBESCO dataset containing 7 emotion classes, 1,440 audio clips from the RAVDESS dataset containing 8 emotion classes, and 8,440 audio clips from the combined dataset of SUBESCO and RAVDESS datasets containing 8 emotion classes. The proposed model achieves recognition accuracy of 96.61% for the SUBESCO dataset, 89.24% for the RAVDESS dataset, and 94.16% for the combined dataset maintaining a training and testing dataset ratio of 80:20 for each case. The performance results of the proposed model show the significant improvements of the state-of-the-art models in achieving accurate and reliable emotion recognition.