Applications for emotion recognition in speech span from mental health evaluation to human-computer interaction. In order to analyze emotional expressions in speech signals, this research introduces a novel method that combines convolutional neural networks (CNNs) (Qayyum et al., Proceedings of the 2019 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), 2019) with attention processes. Speech emotion identification systems have advanced in a number of ways, including the application of deep learning models and novel temporal and auditory variables. This research presents a two-dimensional Convolutional Neural Network (CNN) and long short-term memory (LSTM) (Xie et al., IEEE/ACM Trans. Audio Speech Lang Process 27:1675–1685, 2019) network combination to develop a self-attention-based deep learning model. This work expands on previous research by conducting extensive experiments on various combinations of spectral and rhythmic information in order to determine the features that perform the best for this task. By modeling speech as Mel-spectrograms, we allow CNNs to capture spatial information while also accounting for the temporal dynamics of emotions. Our parallel CNN-Transformer network has an accuracy of 74%, followed by the parallel CNN-BLSTM-Attention at 60%, outperforming standalone models. Notably, our solution requires fewer parameters, increasing efficiency while maintaining performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Emotion Classification Through CNN Models for Speech Analysis

  • Sri Samyu Tankasala,
  • Sai Harshitha Peddi,
  • Sushma Bodpati,
  • Jahnavi Kuddigana,
  • C. P. Prathibhamol

摘要

Applications for emotion recognition in speech span from mental health evaluation to human-computer interaction. In order to analyze emotional expressions in speech signals, this research introduces a novel method that combines convolutional neural networks (CNNs) (Qayyum et al., Proceedings of the 2019 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), 2019) with attention processes. Speech emotion identification systems have advanced in a number of ways, including the application of deep learning models and novel temporal and auditory variables. This research presents a two-dimensional Convolutional Neural Network (CNN) and long short-term memory (LSTM) (Xie et al., IEEE/ACM Trans. Audio Speech Lang Process 27:1675–1685, 2019) network combination to develop a self-attention-based deep learning model. This work expands on previous research by conducting extensive experiments on various combinations of spectral and rhythmic information in order to determine the features that perform the best for this task. By modeling speech as Mel-spectrograms, we allow CNNs to capture spatial information while also accounting for the temporal dynamics of emotions. Our parallel CNN-Transformer network has an accuracy of 74%, followed by the parallel CNN-BLSTM-Attention at 60%, outperforming standalone models. Notably, our solution requires fewer parameters, increasing efficiency while maintaining performance.