Enhancing Speech Emotion Recognition Combining Silence Elimination and Attention Model with a Novel CNN Architecture
摘要
A natural interface for human–computer interaction is automatic speech emotion recognition, but it encounters challenges in handling non-emotional speech segments, especially silence, since it is non-emotional speech segmentation that requires recognition. Using a carefully designed attention model and a silence elimination approach, this paper proposes a novel method of enhancing emotion recognition performance by combining silence elimination with attention. In the present study, an enhanced Convolutional Neural Network architecture has been proposed, comprising of nine layers, utilizing a combination of convolutional 1D As a result of the research, CNNs have been employed in a revolutionary speech-emotion detection system that reaches an impressive 96.00% accuracy rate. There is great potential in this breakthrough to develop social robots and conversational robots capable of conveying nuanced human emotions. This breakthrough has marked a significant improvement in accuracy compared to prior models based on the same dataset. In addition to the broader classifications of positive, negative, and neutral emotions, the model's accuracy is evaluated across a wide range of emotion classes, such as anger, calmness, fear, happiness, and sadness for both genders. There is considerable evidence to suggest that combined noise cancellation and attention models are more effective than individual models of noise cancellation or attention.