A Statistical WavLM Embedding Features with Auto-Encoder for Speech Emotion Recognition
摘要
Speech Emotion Recognition (SER) is an emerging field that encompasses various disciplines such as Human-Computer Interaction (HCI), Natural Language Processing (NLP), computer vision, and cognitive sciences like psychology and social sciences. The primary objective of this SER study is to analyze and quantify human emotions using a combination of statistical feature extraction and Deep Learning (DL) techniques. To achieve this goal, the Mi-Auto-Encoder (MiAE) is proposed to compress the embedding features representation of the WavLM model; in addition, a dense layer is incorporated to classify the different emotions. The SER experiments were conducted on the widely used Interactive Emotional Dyadic Motion Capture (IEMOCAP) English reference database. The results revealed promising performance, with accuracies of \(77.57\%\) and \(76.17\%\) achieved on the validation and test data, respectively. The proposed SER system was evaluated and compared to state-of-the-art studies, demonstrating its effectiveness.