A novel convolution neural network architecture with fully connected network for efficient speech emotion recognition system
摘要
Speech emotion recognition determines the emotion of a speaker utilizing only his/her spoken utterance. This research faces difficulty in achieving high speech emotion recognition accuracy due to differences in speaker’s emotional intensity. This study has proposed a novel 1D CNN architecture integrated with fully connected layers which is able to capture the emotion characteristics in speech efficiently for addressing the above mentioned issue. Experiments conducted on RAVDESS dataset which consists of two levels of speaker emotion intensity, namely, ‘strong’ and ‘normal’ have shown better overall performance for the proposed architecture over the Baseline. The proposed method has achieved 4.1% improvement in average speech emotion recognition accuracy compared to the Baseline across 8 different emotions. Further the proposed method has demonstrated superior performance for recognizing neutral & sad emotions. It has achieved 14 percent and 9 percent higher speech emotion recognition accuracy over the Basline method for neutral and sad emotions, respectively.