错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Experimental Analysis of Emotion Recognition in Voice Using MFCC and Deep Neural Network

  • Monika Khatkar,
  • Asha Sohal,
  • Ramesh Kait

摘要

As the amount of human–computer connection has increased steadily over the past few years and emotion recognition in voice has attracted a lot of attention. This study uses DNN (deep neural networks) and MFCC (mel-frequency cepstral coefficients) to provide a novel method for audio emotion identification. The suggested approach seeks to precisely categorise speakers’ emotional states based on their auditory signals. The features from the voice samples are retrieved using MFCC, and the deep neural network model uses these characteristics as input. To associate the retrieved MFCC characteristics with the appropriate emotional labels, this model was trained using a sizable dataset of speech samples that had been categorised and covered a wide spectrum of emotions. This concept is used to capture discriminative features for voice sentiments recognition and deep neural networks to achieve state-of-the-art performance, outperforming traditional machine learning methods in terms of accuracy, precision, recall, and F1-score with RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song). The outcomes show how well the MFCC features work in extracting emotion-related data from audio signals. Additionally, the DNN obtains an average performance of 61%, demonstrating its capacity to effectively learn and categorise emotions.