Emotion Detection Through Advanced Audio Feature Analysis Using Deep Learning Techniques
摘要
Speech emotion recognition is a growing field that can be utilized in a variety of human-computer interaction applications such as assistive technology, gaming, and virtual reality. This research paper examines SER in depth utilizing deep learning techniques. The research work also extracts features from four separate speech datasets (CREMA-D, RAVDESS, SAVEE, and TESS). The model is able to recognize distinct vocal patterns that accurately represent individual demands. The MFCC features are extremely useful in determining the types of emotion in audios. In comparison to CNN, MobileNet and EfficientNet have lesser accuracy of 92% and 79%, respectively. The test findings show that there is a significant improvement in precision when compared to the current standard models.