Speech emotion recognition is considered one of the important topics in Human-Computer Interaction (HCI). It is considered as the task of identifying classes of emotions by using effective features from inputs speech. In the last few years, speech emotion recognition has been considered one of the most essential and most challenging problems. The challenges came from: (1) Uncertainty about the effective characteristics of speech signals. (2) The available collections have a few recordings that have been collected by a few speakers. (3) The lengthy time required to annotate with emotional labels. In this paper, a practical solution to train Dense Neural Networks (DNNs), Residual Networks (RN) and Multiscale Convolution Neural Networks (MCNN) for emotion-speech-emotion recognition by using a very a limited number of samples which selected by active learning (AL) is proposed. The Greedy method (GM), depending on the record’s feature, is used to select samples from a pool of samples. The goal of the proposed model is to predict concordance correlation coefficients of arousal and valence in the regression process. The results showed that the use of active learning-based DNNs and RN led to better performance with limited training data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid Intelligent Model for Speech Emotion Recognition Using Active Learning and Residual Network

  • Ghada Dahy,
  • Aboul Ella Hassanien,
  • Sameh H. Basha

摘要

Speech emotion recognition is considered one of the important topics in Human-Computer Interaction (HCI). It is considered as the task of identifying classes of emotions by using effective features from inputs speech. In the last few years, speech emotion recognition has been considered one of the most essential and most challenging problems. The challenges came from: (1) Uncertainty about the effective characteristics of speech signals. (2) The available collections have a few recordings that have been collected by a few speakers. (3) The lengthy time required to annotate with emotional labels. In this paper, a practical solution to train Dense Neural Networks (DNNs), Residual Networks (RN) and Multiscale Convolution Neural Networks (MCNN) for emotion-speech-emotion recognition by using a very a limited number of samples which selected by active learning (AL) is proposed. The Greedy method (GM), depending on the record’s feature, is used to select samples from a pool of samples. The goal of the proposed model is to predict concordance correlation coefficients of arousal and valence in the regression process. The results showed that the use of active learning-based DNNs and RN led to better performance with limited training data.