<p>Speech emotion recognition determines the emotion of a speaker utilizing only his/her spoken utterance. This research faces difficulty in achieving high speech emotion recognition accuracy due to differences in speaker’s emotional intensity. This study has proposed a novel 1D CNN architecture integrated with fully connected layers which is able to capture the emotion characteristics in speech efficiently for addressing the above mentioned issue. Experiments conducted on RAVDESS dataset which consists of two levels of speaker emotion intensity, namely, ‘strong’ and ‘normal’ have shown better overall performance for the proposed architecture over the Baseline. The proposed method has achieved 4.1% improvement in average speech emotion recognition accuracy compared to the Baseline across 8 different emotions. Further the proposed method has demonstrated superior performance for recognizing neutral &amp; sad emotions. It has achieved 14 percent and 9 percent higher speech emotion recognition accuracy over the Basline method for neutral and sad emotions, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel convolution neural network architecture with fully connected network for efficient speech emotion recognition system

  • Vandana Singh,
  • Swati Prasad

摘要

Speech emotion recognition determines the emotion of a speaker utilizing only his/her spoken utterance. This research faces difficulty in achieving high speech emotion recognition accuracy due to differences in speaker’s emotional intensity. This study has proposed a novel 1D CNN architecture integrated with fully connected layers which is able to capture the emotion characteristics in speech efficiently for addressing the above mentioned issue. Experiments conducted on RAVDESS dataset which consists of two levels of speaker emotion intensity, namely, ‘strong’ and ‘normal’ have shown better overall performance for the proposed architecture over the Baseline. The proposed method has achieved 4.1% improvement in average speech emotion recognition accuracy compared to the Baseline across 8 different emotions. Further the proposed method has demonstrated superior performance for recognizing neutral & sad emotions. It has achieved 14 percent and 9 percent higher speech emotion recognition accuracy over the Basline method for neutral and sad emotions, respectively.