Recently, there has been interest in classifying emotions using audio inputs and machine learning methods. Because a single statement might be delivered in a variety of emotional circumstances, textual data alone is insufficient for identifying the emotions, necessitating the adoption of novel feature extraction approaches. This paper aims to provide an improved framework which uses a detailed study on audio sound features like mel-frequency cepstral coefficients (MFCCs), zero crossing rate (ZCR), spectral bandwidth and many other features for audio analysis and uses the same for predicting emotions. The paper also emphasises on various audio augmentation techniques to improve the generalising power of the model. With the above-mentioned improvements in the feature vector and a carefully selected deep neural network, state-of-the-art classification accuracy is achieved on the two state-of-the-art datasets Toronto emotional speech set (TESS) and the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). The paper is concluded with the improved approach for emotion recognition using audio signals and leads to future scope of efficient speech recognition applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Emotion Recognition in Audio Signals: Leveraging Novel Features and Deep Learning for Improved Classification

  • Poonam Chaudhary,
  • Neeraj Choudhary

摘要

Recently, there has been interest in classifying emotions using audio inputs and machine learning methods. Because a single statement might be delivered in a variety of emotional circumstances, textual data alone is insufficient for identifying the emotions, necessitating the adoption of novel feature extraction approaches. This paper aims to provide an improved framework which uses a detailed study on audio sound features like mel-frequency cepstral coefficients (MFCCs), zero crossing rate (ZCR), spectral bandwidth and many other features for audio analysis and uses the same for predicting emotions. The paper also emphasises on various audio augmentation techniques to improve the generalising power of the model. With the above-mentioned improvements in the feature vector and a carefully selected deep neural network, state-of-the-art classification accuracy is achieved on the two state-of-the-art datasets Toronto emotional speech set (TESS) and the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). The paper is concluded with the improved approach for emotion recognition using audio signals and leads to future scope of efficient speech recognition applications.