Speech emotion recognition (SER) is an important technology that improves, sentiment analysis, mental health monitoring, and human machine collaboration. Existing SER approaches struggle to correctly capture the complex, nonlinear, and temporal relationships seen in speech signals. To overcome these challenges, we offer a new technique based on a CNN-LSTM hybrid model optimized via Bayesian optimization. The CNN component successfully recovers spatial characteristics from raw audio data, whereas the LSTM component collects temporal relationships and contextual information. Bayesian optimization is used to fine-tune hyperparameters, which improves model performance and resilience. The experimental findings show considerable increases in accuracy and other performance indicators when compared to typical SER models. Here we are using audio datasets Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) on which we obtained an accuracy of 95% and Crowd Sourced Emotional Multimodal Actors Dataset (CREMA-D) with an accuracy of 89%, after applying the CNN-LSTM model and then subsequently Bayesian optimization. We also employ confusion matrix and classification report to further evaluate the performance. Finally, our work enhances the area of SER by providing a more accurate and efficient model that has potential applications in a multitude of disciplines, including customer support, market analysis, entertainment, and mental wellness screening. Future studies will look into further improvements and real-world applications of the proposed framework.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced Speech Emotion Recognition: A Hybrid Deep Learning Model and Bayesian Optimization

  • Deepti Gupta,
  • Bhawna Jain,
  • Arun Sharma

摘要

Speech emotion recognition (SER) is an important technology that improves, sentiment analysis, mental health monitoring, and human machine collaboration. Existing SER approaches struggle to correctly capture the complex, nonlinear, and temporal relationships seen in speech signals. To overcome these challenges, we offer a new technique based on a CNN-LSTM hybrid model optimized via Bayesian optimization. The CNN component successfully recovers spatial characteristics from raw audio data, whereas the LSTM component collects temporal relationships and contextual information. Bayesian optimization is used to fine-tune hyperparameters, which improves model performance and resilience. The experimental findings show considerable increases in accuracy and other performance indicators when compared to typical SER models. Here we are using audio datasets Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) on which we obtained an accuracy of 95% and Crowd Sourced Emotional Multimodal Actors Dataset (CREMA-D) with an accuracy of 89%, after applying the CNN-LSTM model and then subsequently Bayesian optimization. We also employ confusion matrix and classification report to further evaluate the performance. Finally, our work enhances the area of SER by providing a more accurate and efficient model that has potential applications in a multitude of disciplines, including customer support, market analysis, entertainment, and mental wellness screening. Future studies will look into further improvements and real-world applications of the proposed framework.