Empowering Robust Speech Emotion Recognition Using Deep Neural Network
摘要
Speech Emotion Recognition, also known as SER, refers to the act of recognizing emotion from a speech input of a human. This is based on the general fact that the emotion of a person is reflected in the tone and pitch of his/her voice. In traditional recognition techniques, handcrafted features may not be capture all relevant information present in the speech signal, leading to suboptimal performance. The paper, proposes a Deep Neural Network based Speech Recognizing technique called DeNSER. The work is aimed at offering advantages like end-to-end learning, automatic feature learning, capturing complex relationships, scalability, and adaptability. The proposed system has used the RAVDESS dataset for training and testing speech and song samples. The implemented DNN-based SER, DeNSER has performed well in recognizing emotions with 89.6% accuracy.